Direct answer: A current AI voice phishing study accepted for Expert Systems with Applications and available on arXiv tested 4,100 U.S. adults and 12 qualitative interviews against AI-generated and human scam scenarios. The authors report up to 36% self-reported compliance in a relative-in-distress scenario and 16.5% overall compliance across five scam categories, while finding that caller persuasiveness was the strongest predictor of compliance and that some AI voice models achieved near-human persuasiveness ratings. Voice-agent buyers should treat this as a production proof gate: validate caller identity, disclosure, transaction limits, fraud monitoring, human escalation, and recovery evidence before voice automation can change records, trigger payments, or collect sensitive information.
What happened
- The paper 'Evaluating AI Models' Capability to Automate Voice Phishing Attacks' was posted to arXiv on July 10, 2026 and is listed by ScienceDirect for Expert Systems with Applications.
- The study reports a large-scale survey experiment with 4,100 U.S. adults and 12 qualitative interviews.
- Participants were exposed to audio recordings or transcripts generated using models including Llama Full Duplex, Sesame, Gemini, OAI AVM, Play.AI, and ElevenLabs, alongside human baselines.
- The authors report up to 36% self-reported compliance in the relative-in-distress category and 16.5% overall compliance across five scam categories.
- BSides Las Vegas is surfacing the work for security practitioners, and CMS has separately warned that AI makes Medicare vishing scams more dangerous by making social engineering more sophisticated.
Why this is trending
- The study moves AI voice fraud from anecdote to measurable procurement risk because it compares multiple voice models, human baselines, text baselines, and scam categories at population scale.
- The headline risk is not that every AI voice is superhuman. The paper argues the main danger is automation economics: low per-call success rates can still become damaging when synthetic calls scale cheaply.
- Voice-agent buyers are simultaneously adopting inbound and outbound AI callers, which means the same capabilities that improve service can also raise verification, disclosure, and recovery expectations.
The Voice Agent Index take
A voice-agent buyer should not approve production voice automation until the fraud packet is visible. The buyer needs caller identity proof, known-good callback paths, AI identity disclosure, consent handling, transaction authority limits, high-risk phrase monitoring, human escalation, fraud-team handoff, audit logs, and customer recovery evidence.
AI Vishing Proof Packet
A buyer checklist for validating voice-agent fraud controls across caller identity, synthetic-voice disclosure, transaction authority, high-risk prompts, monitoring, human escalation, and customer recovery evidence.
| Proof item | Why it matters | Buyer ask |
|---|---|---|
| Caller identity | AI vishing succeeds when a caller sounds persuasive enough to move a person before identity is verified. | Require known-good callback workflows, account verification rules, caller-intent checks, spoofing controls, and proof that agents cannot rely on voice alone. |
| Synthetic-voice disclosure | Customers, patients, and policyholders need to know when they are interacting with an AI voice agent, especially in regulated or sensitive workflows. | Ask for disclosure wording, consent timing, recording rules, language variants, opt-out handling, and logs showing disclosure happened. |
| Transaction authority | A convincing voice should not be enough to move money, change account access, expose records, reset credentials, or trigger high-risk actions. | Define which actions the voice agent may never complete without step-up verification, human approval, or a known-good channel. |
| High-risk prompt detection | Fraud patterns include urgency, relatives in distress, Medicare or insurance pressure, password resets, payment requests, and sensitive-data collection. | Test detection and escalation for urgency, distress, payment, credentials, health data, identity documents, unusual callbacks, and account takeover language. |
| Human escalation | Customers need a protected path to a person when the voice agent detects fraud, uncertainty, distress, or unsupported requests. | Require supervisor routing, fraud-team handoff, callback SLA, agent authority, and evidence from test calls with risky scenarios. |
| Recovery evidence | Fraud controls are incomplete unless the organization can show what happened after a bad call, suspicious request, or mistaken disclosure. | Keep audit logs, recordings where lawful, transcript exports, fraud tags, customer notices, account locks, reversal steps, and incident closure proof. |
What buyers should do next
- List every production or pilot voice-agent workflow that can collect sensitive information, update records, reset access, schedule callbacks, or influence payments.
- Define no-go actions that require human approval, step-up verification, or a known-good callback channel.
- Build a vishing test set with relative-in-distress, payment, credential, Medicare or insurance, urgent callback, and spoofed-identity scenarios.
- Run disclosure, consent, fraud-detection, escalation, and recovery tests through the exact telephony path planned for production.
- Add fraud outcomes, suspicious-call tags, human handoff quality, customer recovery, and rollback thresholds to the weekly voice-agent QA report.
Turn this brief into a vendor packet
Make the vendor prove the workflow before the demo gets polished.
Use the RFP generator and call-test script to turn this news framework into concrete evidence requests, acceptance tests, and escalation rules for your own voice AI rollout.
Buyer FAQs
What did the AI voice phishing study find?
The study reports that in a survey experiment with 4,100 U.S. adults, up to 36% of participants said they would or might comply in a relative-in-distress scenario, and overall compliance across five scam categories was 16.5%.
Does the study say AI voices are always more persuasive than humans?
No. The authors emphasize automation economics and caller persuasiveness. Some models achieved near-human ratings, but the main operational risk is that scalable synthetic calling can make even modest compliance rates dangerous.
What proof should voice-agent buyers request?
Ask for caller identity checks, AI disclosure logs, consent handling, transaction authority limits, high-risk prompt detection, human escalation, fraud-team handoff, audit logs, and customer recovery evidence.
Sources
- ScienceDirect: Expert Systems with Applications article page for 'Evaluating AI Models' Capability to Automate Voice Phishing Attacks.'
- arXiv: July 10, 2026 preprint with the abstract, authors, model list, 4,100-participant survey, 12 interviews, compliance rates, and automation-economics framing.
- BSides Las Vegas: Security-conference listing surfacing the vishing study, model list, scam scenarios, up-to-36% success signal, and detection focus.
- CMS: February 11, 2026 CMS security guidance explaining why AI makes Medicare vishing scams more sophisticated and dangerous.