arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.24127cs.CRcs.CLcs.CYcs.LG

诈骗电话剖析:10000个真实诈骗与垃圾电话揭示电话诈骗者的运作方式

Anatomy of a Scam Call: What 10,000 real scam and spam calls reveal about how phone scammers operate

Ethan Traister, Ankit Raj, Jiaqi Gan, Xingyu Shen, Tyler Wu, Yuchen Zhou, Tommy Duong, Kidus Zewde, Siying Chen, Simiao Ren

中文总结 AI 辅助

该研究通过分析10211个真实电话数据,揭示诈骗电话遵循办公时间规律,会针对目标年龄调整施压程度但固定索要内容,且仅从开场话术即可早期检测,词袋分类器表现与微调模型相当。

中文摘要 AI 辅助

电话诈骗无处不在且代价高昂,但其内部运作机制很少被大规模观测到。我们分析了一个完整的语料库,包含10211个打入的诈骗和垃圾电话——共913小时音频,来自5780个不同号码的330956段转录对话轮次,这些数据由一个AI语音代理蜜罐在54天内收集,该蜜罐会接听来电并让通话者持续交谈,相关内容已在配套数据描述符中介绍。我们将直接诈骗(即套取敏感信息)与规模更大但合法的掠夺性线索生成(“垃圾邮件”)区分开来,后者为直接诈骗提供支持。诈骗运营遵循办公时间规律:工作日的通话量是周末的6.6倍;数千个一次性号码使用少量重复的话术脚本(30个开场聚类,前5个聚类承载了一半的流量);来电者获取身份锚点(家庭地址和出生日期)的频率远高于支付凭证,其手段是坚持不懈和制造权威,而非明显的威胁。我们的核心实验探究:谁接听电话是否会产生影响?每个植入的线索都带有10个随机均匀抽取的虚构身份之一,因此诈骗运营接触到的身份在来电者出现前就已确定。在1823个随机通话中,诈骗者针对目标表面年龄每增加十年,会多花费约15%的对话轮次(比率为1.15,95%置信区间1.08-1.23;随机检验p值为0.005)——但他们索要的内容并未改变(26.3%的通话会提出敏感信息请求;每十年的赔率比为0.99,95%置信区间0.90-1.08)。第二个实验将早期检测作为基准:仅从诈骗者的开场话术来看,在来电者不重叠的划分下,从第一句就能以0.72的ROC-AUC值预测升级,到第八句时达到0.87,且简单的词袋分类器与微调后的设备端语言模型表现相当。电话诈骗是一个模板化行业,其差异在于对目标的施压程度,而非想要的东西。

英文摘要

Telephone fraud is pervasive and costly, but its inner workings are rarely observed at scale. We analyze a complete corpus of 10,211 inbound scam and spam calls -- 913 hours of audio and 330,956 transcribed turns from 5,780 distinct numbers -- collected over 54 days by an AI voice-agent honeypot that answered callers and kept them talking, and introduced in a companion data descriptor. We separate outright scams, which solicit sensitive information, from the larger stream of predatory but legal lead generation ("spam") that feeds them. Scam operations keep office hours (6.6x more calls per weekday than weekend day); thousands of disposable numbers run a small catalog of recycled scripts (thirty opening clusters, half the traffic in the top five); and callers solicit identity anchors -- a home address and a date of birth -- far more often than payment credentials, pressing through persistence and manufactured authority rather than overt threats. Our central experiment asks: does it matter who picks up? Every seeded lead carried one of ten fictitious identities drawn uniformly at random, so the identity a fraud operation reaches is fixed before the caller exists. Across 1,823 randomized calls, scammers spent about 15% more conversational turns per decade of the target's apparent age (rate ratio 1.15, 95% CI 1.08-1.23; randomization p = 0.005) -- yet what they asked for did not change (26.3% of calls reached a request for sensitive information; odds ratio 0.99 per decade, 95% CI 0.90-1.08). A second experiment casts early detection as a benchmark: from a scammer's opening lines alone, on a caller-disjoint split, escalation is predictable at 0.72 ROC-AUC from the first line and 0.87 by the eighth, and a plain bag-of-words classifier matches a fine-tuned on-device language model. Telephone fraud emerges as a templated industry that varies how hard it works a target, but not what it wants.

补充信息

↑