AI Voice Phishing Performs on Par With Human Scammers at a Fraction of the Cost
Fred Heiding
SubscribeSign in
AI Voice Phishing Performs on Par With Human Scammers at a Fraction of the Cost
Fred Heiding and Simon Lermen<br>Jul 20, 2026
Share
TL;DR: We ran a large-scale human-subject study (n=4,100) to measure susceptibility to AI-powered voice phishing, using six leading AI voice models. They achieved high compliance rates, with up to 36% of participants who would or might fall for the scam. Participants struggled to distinguish AI-powered voices from human callers.
Thanks for reading! Subscribe for free to receive new posts and support my work.
Subscribe
Full paper: https://arxiv.org/abs/2607.09970
This post is intended to be a brief summary of the main findings, which include:<br>Models like Sesame and ElevenLabs performed on par with humans in many experiments, and sometimes even outperformed them.
Our economic analysis suggests that AI voice phishing is already profitable for several of these models.
Caller persuasiveness was the strongest predictor of compliance, regardless of whether the caller was perceived as AI or human.
Participants who frequently used AI systems were no better at identifying AI-powered voices than those with no AI exposure.
Abstract
Voice phishing (vishing) attacks have traditionally been limited by the need for human operators. The rapid emergence of high-quality AI voice synthesis and large language models (LLMs) reduces this bottleneck and enables scalable, automated scams. In this paper, we conduct a large-scale survey experiment (N=4100) and qualitative interviews (N=12) to assess U.S. adults’ susceptibility to AI-powered voice phishing attacks. Participants were exposed to audio recordings or transcripts of scam scenarios generated using leading voice models such as Llama Full Duplex (Llama FD), Sesame, Gemini, OAI AVM, Play.AI, and ElevenLabs and the corresponding human baselines. The results show high compliance rates. Up to 36% of participants would or might comply with phishing requests in the “relative-in-distress” category. Overall compliance rate across all five scam categories was 16.5%, a striking figure given the low cost and high scalability of AI-automated voice phishing. Caller persuasiveness was the strongest predictor of compliance and certain models (most notably Sesame) achieved ratings comparable to human voices, or sometimes even slightly surpassing them. Our economic analysis suggests that while human-operated vishing is unprofitable at US wages, AI-powered vishing appears to be economically viable for several models. The primary risk of present-day AI-enabled vishing thus lies in the economics of automation rather than novel or “superhuman” persuasive techniques, though these cannot be ruled out for future systems. This raises significant concerns for the design of AI systems, consumer protection, and model release policies.<br>Method
In a brief summary, the method consists of the following steps:<br>Recruited 4,100 participants representative of U.S. adult internet users.
Comparing six AI voice models, participants were randomly assigned one of 37 experimental conditions based on five scam scenarios:
1. MasterCard scam<br>2. Gmail scam<br>3. Donation scam<br>4. Police-Grandma scam<br>5. Sister-in-distress scam<br>Each participant evaluated one audio recording or transcript of a randomly assigned scenario using five-point scales:
1. Caller sentiment<br>2. Persuasiveness<br>3. Trustworthiness<br>4. Human-likeness<br>5. Compliance with scam<br>Statistical analysis was conducted for significance among continuous outcomes, comparison of AI models to human voice control, willingness to comply, and predictor variables.
12 qualitative interviews were conducted to complement quantitative findings with deeper insights into participant reasoning and perceptions.
Results
This section presents findings from our large-scale evaluation of AI-powered voice phishing. We organize our results to address our primary research questions systematically, examining (1) how AI models perform in neutral contexts, (2) how scam content affects perception, and (3) what factors drive susceptibility to AI-powered attacks.<br>How AI models perform in neutral contexts
To establish baseline differences between AI models independent of scam context, we compared participant responses to neutral (non-scam) scenarios. Sesame was rated as significantly more humanlike than Llama FD, OpenAI AVM, and Gemini, but did not differ from Play.AI or ElevenLabs. When compared against an authentic human voice, Sesame was the only model that achieved statistical parity on human-likeness. The ElevenLabs cloned voice also matched human-level performance, suggesting that high-quality voice cloning can achieve human-level naturalness in neutral contexts.<br>How scam content affects perception
The transition from neutral to scam scenarios produced negative effects across all measured dimensions.
Impact of Scam Context on AI Model...