[2607.28165] Piggybacking on Perception: Stealthy Concurrent Audio Prompt Injections against Multimodal LLM Agents
Skip to main content
Search arXiv
Press Enter to search · Advanced search
-->
Computer Science > Cryptography and Security
arXiv:2607.28165 (cs)
[Submitted on 30 Jul 2026]
Title:Piggybacking on Perception: Stealthy Concurrent Audio Prompt Injections against Multimodal LLM Agents
Authors:Mingxiao Liu (1), Yitong Li (1), Haoren Zhao (1), Yaoxiang Bian (1), Jianan Ma (1 and 2), Jian Zhang (1), Jialuo Chen (3 and 2), Xinhao Deng (4 and 2), Zhen Wang (1) ((1) Hangzhou Dianzi University, (2) Ant Group, (3) Zhejiang University, (4) Tsinghua University)<br>View a PDF of the paper titled Piggybacking on Perception: Stealthy Concurrent Audio Prompt Injections against Multimodal LLM Agents, by Mingxiao Liu (1) and 11 other authors
View PDF<br>HTML (experimental)
Abstract:Large Language Model (LLM)-driven multimodal agents are increasingly deployed to execute autonomous tasks via continuous audio interaction. While this paradigm enhances interaction naturalness, it introduces a critical yet under-explored attack surface, as audio inputs inevitably contain environmental noise beyond user control. In this paper, we investigate concurrent audio prompt injection attacks targeting multimodal agents. Distinct from traditional acoustic attacks on voice devices, we propose novel techniques for instruction augmentation and scenario concealment. These methods allow malicious audio instructions to imperceptibly "piggyback" onto user speech, thereby hijacking agents to execute malicious actions. To systematically quantify this threat, we construct AudioAgentSecurity, the first comprehensive benchmark for audio instruction injection attacks, encompassing 8 real-world task scenarios and 10 distinct attack patterns. We evaluate 11 state-of-the-art agents, including Gemini 3 Pro and GPT-4o-audio. Notably, our methods achieve an average Attack Success Rate (ASR) of 69.10\% against the advanced Gemini 3 Pro. To counter this threat, we further introduce Cascaded Audio Decoupling and Verification (CADV), a defense mechanism based on source separation and consistency analysis. Compared with existing prompt-level defenses, CADV leverages acoustic source separation and cross-modal consistency analysis to detect audio instruction injections more robustly, achieving over 90\% detection success across diverse attack vectors. Finally, real-world experiments with human volunteers on Doubao AI Smartphone in diverse dynamic real-world scenarios confirm the attacks' high stealth and efficacy, while demonstrating that our defense reliably mitigates these vulnerabilities.
Comments:<br>17 pages, 8 figures, The code is publicly available at this https URL
Subjects:
Cryptography and Security (cs.CR)
Cite as:<br>arXiv:2607.28165 [cs.CR]
(or<br>arXiv:2607.28165v1 [cs.CR] for this version)
https://doi.org/10.48550/arXiv.2607.28165
Focus to learn more
arXiv-issued DOI via DataCite (pending registration)
Submission history<br>From: Mingxiao Liu [view email]<br>[v1]<br>Thu, 30 Jul 2026 13:04:15 UTC (13,293 KB)
Full-text links:<br>Access Paper:
View a PDF of the paper titled Piggybacking on Perception: Stealthy Concurrent Audio Prompt Injections against Multimodal LLM Agents, by Mingxiao Liu (1) and 11 other authors<br>View PDF<br>HTML (experimental)<br>TeX Source
view license
Current browse context:
cs.CR
next >
new<br>recent<br>| 2026-07
Change to browse by:
cs
References & Citations
NASA ADS<br>Google Scholar
Semantic Scholar
export BibTeX citation<br>Loading...
BibTeX formatted citation
×
loading...
Data provided by:
Bookmark
Bibliographic Tools
Bibliographic and Citation Tools
Bibliographic Explorer Toggle
Bibliographic Explorer (What is the Explorer?)
Connected Papers Toggle
Connected Papers (What is Connected Papers?)
Litmaps Toggle
Litmaps (What is Litmaps?)
scite.ai Toggle
scite Smart Citations (What are Smart Citations?)
Code, Data, Media
Code, Data and Media Associated with this Article
alphaXiv Toggle
alphaXiv (What is alphaXiv?)
Links to Code Toggle
CatalyzeX Code Finder for Papers (What is CatalyzeX?)
DagsHub Toggle
DagsHub (What is DagsHub?)
GotitPub Toggle
Gotit.pub (What is GotitPub?)
Huggingface Toggle
Hugging Face (What is Huggingface?)
ScienceCast Toggle
ScienceCast (What is ScienceCast?)
Demos
Demos
Replicate Toggle
Replicate (What is Replicate?)
Spaces Toggle
Hugging Face Spaces (What is Spaces?)
Spaces Toggle
TXYZ.AI (What is TXYZ.AI?)
Related Papers
Recommenders and Search Tools
Link to Influence Flower
Influence Flower (What are Influence Flowers?)
Core recommender toggle
CORE Recommender (What is CORE?)
Author
Venue
Institution
Topic
About arXivLabs
arXivLabs: experimental projects with community collaborators
arXivLabs is a framework that allows collaborators to develop and share new arXiv features directly on our website.
Both individuals and organizations that work with...