Harvard/Meta Study Quantifies AI Voice-Phishing (Vishing) Persuasiveness Gap: 16.5% Compliance, 70.3% Detection, and Profitable Economics for Attackers — Threadlinqs Intelligence
As of 2026-07-17, Harvard/Meta Study Quantifies AI Voice-Phishing (Vishing) Persuasiveness Gap: 16.5% Compliance, 70.3% Detection, and Profitable Economics for Attackers is a medium-severity threat intel threat, tracked by Threadlinqs Intelligence with 9 detection rules (Splunk SPL, Microsoft KQL, Sigma) and 15 indicators of compromise.
Threat ID: TL-2026-1431 · Severity: MEDIUM · Status: ACTIVE · Category: THREAT_INTEL
A peer-reviewed study by Harvard Kennedy School, Harvard SEAS, and Meta researchers (arXiv:2607.09970, accepted to Expert Systems with Applications) surveyed 4,100 US adults against recordings from
Researchers Fred Heiding, Claudio Mayrink Verdun, Simon Lermen, Andrew Kao, Vitor Albiero, Lauren Deason, Irina-Elena Veliche, and Christine Lehane (affiliations: Harvard Kennedy School, Harvard School of Engineering and Applied Sciences, Meta Platforms Inc., and independent researchers) published 'Evaluating AI Models' Capability to Automate Voice Phishing Attacks' (arXiv:2607.09970), accepted for publication in Expert Systems with Applications and covered by Help Net Security on 2026-07-17.
The study combined a large-scale survey experiment (N=4,100 US adults) with qualitative interviews (N=12). Participants were exposed to audio recordings or transcripts of five vishing scenarios generated by six leading voice-AI systems — Llama Full-Duplex (Meta, experimental mode, April 2025), OpenAI Advanced Voice Mode ('Sol' voice, GPT-4o), Google Gemini ('Ursa' voice, mobile app), Sesame ('Maya' voice, open-source conversational model), Play.AI ('Celeste' voice), and ElevenLabs (voice cloning) — alongside human-caller control recordings, with model refusals, disclaimers, and generation beeps edited out to simulate a competent attacker.
The five scam scenarios were: (1) MasterCard credential phishing impersonating card support (~7-15% compliance), (2) Gmail credential phishing impersonating Google support (~6-13% compliance), (3) a charitable Donation solicitation from an unknown caller (20-32% compliance), (4) a 'Police-Grandma' bail/relative-in-trouble scenario from an unknown caller (24-25% compliance), and (5) a 'Sister-in-Distress' scenario using an ElevenLabs-cloned voice claiming a family emergency requiring immediate funds, which produced the highest compliance at 36.1% (combined 'yes'/'maybe' responses; 6.5% hard 'yes'). Overall compliance across all five scenarios and both AI and human callers averaged 16.5%.
Detection performance varied sharply by model: participants correctly flagged OpenAI AVM as synthetic 93.7% of the time, Llama FD 88.6%, Gemini 88.4%, Play.AI 85.6%, ElevenLabs clones 75.8%, but Sesame only 66.3% — the hardest system to distinguish from a human. Overall voice-AI detection accuracy across all six systems was 70.3%, while real human callers were incorrectly flagged as synthetic between 54.2% and 67.4% of the time, indicating participants applied a strong 'guess AI' bias rather than genuine discrimination ability. Neither self-reported AI familiarity nor voice-assistant usage improved detection (51.2% for frequent AI users vs. 54.4% for non-users, not statistically significant).
Regression analysis identified message persuasiveness as the dominant predictor of victim compliance (odds ratio 2.58, p<.001), followed by caller sentiment (OR=1.64) and a 9-item trustworthiness scale covering expertise, credibility, empathy, urgency, and confidence (OR=1.48, Cronbach's alpha=.927). A separate 6-item Partner Modelling Questionnaire human-likeness scale (alpha=.93) did not independently predict compliance once persuasiveness and trustworthiness were controlled for — meaning voice realism itself was not the primary driver of victim harm; script quality and emotional manipulation were. Participants who cited 'repetitive responses' (OR=1.38, p=.048), 'long-winded responses' (OR=3.80, p<.001), or 'unnatural phrasing' in text transcripts (OR=3.39, p<.001) were more likely to correctly detect AI, but these cues went unnoticed by the majority of participants despite being present in the recordings.
The paper's economic model computed expected hourly attacker profit per voice system using a revenue formula r_j(t,X_i) = m(X_i) x p_j(t,X_i) x q, where m(X_i) is average extractable value per victim (~$450), p_j is per-call success probability, and q is a 0.6% call-to-contact conversion rate calibrated from marketing-industry data. Human operators paid US minimum-adjacent wages ($34.55/hr) were unprofitable at -$27.10/hr expected profit despite a 23% success rate, because human labor cost dominates. Among AI systems, Gemini
Target sectors: financial services, telecoms, consumer general public, government administration, elderly vulnerable populations
Target regions: North America, united states of america
Detections & IOCs
As of 2026-07-28, this threat has 9 detection rule(s) across Splunk SPL, Microsoft KQL and Sigma, and 15 indicator(s) of compromise. Detection query text and full IOC values are available to authenticated users and programmatically via the Threadlinqs MCP server (Purple tier). View plans.
THREAT_INTEL, MEDIUM, threat intelligence, cybersecurity, T1589, T1593.001, T1587.001, T1585, T1583.006, T1566.004, T1199, T1204.001, T1621, T1539