arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

评估人工智能模型自动实施语音网络钓鱼攻击的能力

Evaluating AI Models' Capability to Automate Voice Phishing Attacks

Fred Heiding, Claudio Mayrink Verdun, Simon Lermen, Andrew Kao, Vitor Albiero, Lauren Deason, Irina-Elena Veliche, Christine Lehane

arXiv 2607.09970首次发表:更新:

AI 中文总结

研究评估美国成年人对人工智能驱动语音网络钓鱼攻击的易感性,通过大规模调查实验和定性访谈,让参与者接触多种语音模型及人类基线生成的诈骗场景,发现高合规率,分析得出人工智能驱动的vishing在经济上可行,其风险在于自动化经济性,引发相关政策担忧。

AI 中文摘要

传统上,语音网络钓鱼(vishing)攻击受限于对人工操作员的需求。高质量人工智能语音合成和大语言模型的迅速出现减少了这一瓶颈,使可扩展的自动诈骗成为可能。本文进行了大规模调查实验(N = 4100)和定性访谈(N = 12),以评估美国成年人对人工智能驱动的语音网络钓鱼攻击的易感性。参与者接触了使用诸如Llama Full Duplex(Llama FD)、Sesame、Gemini、OAI AVM、this http URL和ElevenLabs等领先语音模型以及相应人类基线生成的诈骗场景的音频记录或文字记录。结果显示出高合规率。在“相对遇险”类别中,高达36%的参与者会或可能遵守网络钓鱼请求。所有五个诈骗类别的总体合规率为16.5%,鉴于人工智能自动语音网络钓鱼的低成本和高可扩展性,这一数字令人震惊。来电者的说服力是合规性的最强预测因素。某些模型(最显著的是Sesame)获得了与人类声音相当的评级,有时甚至略高于人类声音。我们的经济分析表明,虽然在美国工资水平下人工操作的vishing无利可图,但人工智能驱动的vishing对几种模型来说在经济上似乎是可行的。当今人工智能支持的vishing的主要风险在于自动化的经济性,而非新颖或“超人”的说服技巧,不过未来系统不能排除这些技巧。这对人工智能系统设计、消费者保护和模型发布政策提出了重大担忧。

英文摘要

Voice phishing (vishing) attacks have traditionally been limited by the need for human operators. The rapid emergence of high-quality AI voice synthesis and large language models (LLMs) reduces this bottleneck and enables scalable, automated scams. In this paper, we conduct a large-scale survey experiment (N=4100) and qualitative interviews (N=12) to assess U.S. adults' susceptibility to AI-powered voice phishing attacks. Participants were exposed to audio recordings or transcripts of scam scenarios generated using leading voice models such as Llama Full Duplex (Llama FD), Sesame, Gemini, OAI AVM, Play$.$AI, and ElevenLabs and the corresponding human baselines. The results show high compliance rates. Up to 36% of participants would or might comply with phishing requests in the "relative-in-distress" category. Overall compliance rate across all five scam categories was 16.5%, a striking figure given the low cost and high scalability of AI-automated voice phishing. Caller persuasiveness was the strongest predictor of compliance and certain models (most notably Sesame) achieved ratings comparable to human voices, or sometimes even slightly surpassing them. Our economic analysis suggests that while human-operated vishing is unprofitable at US wages, AI-powered vishing appears to be economically viable for several models. The primary risk of present-day AI-enabled vishing thus lies in the economics of automation rather than novel or "superhuman" persuasive techniques, though these cannot be ruled out for future systems. This raises significant concerns for the design of AI systems, consumer protection, and model release policies.

CommentsUpdated to the published version. Published in Expert Systems with Applications, Volume 332, Part D

DOI:10.1016/j.eswa.2026.133620

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑