arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

来自主动语音代理蜜罐的真实诈骗和垃圾电话对话语料库

A Corpus of Real Scam- and Spam-Call Conversations from an Active Voice-Agent Honeypot

Ethan Traister, Dennis Tsang Ng, Siyu Zhang, Huaiyu Guo, Tommy Duong, Tyler Wu, Yuchen Zhou, Xingyu Shen, Jiaqi Wu, Simiao Ren

arXiv 2609.29528首次发表:更新:

AI 中文总结

本研究通过主动语音代理蜜罐收集真实诈骗电话对话数据集,包含53天内10,015通电话及逐轮转录和标签,并验证了其真实性,发现基于合成数据的检测方法在真实流量上精度大幅下降。

AI 中文摘要

欺诈者与其目标之间的真实对话是研究电话诈骗最有信息量的素材之一,但也是最稀缺的:被动蜜罐绝大多数捕获的是自动消息和挂断,大规模研究描述的是呼叫元数据而非对话,而人工诱骗诈骗无法规模化。我们提出了一个由主动语音代理蜜罐收集的真实诈骗电话对话数据集。专用号码被植入欺诈操作所利用的潜在客户生成渠道;来电者由低延迟对话代理接听,该代理采用合理的目标人物角色并维持互动,同时每通电话都被录音、转录并自动标记。在最初的53天窗口内,我们捕获了10,015通入站诈骗和垃圾电话(其中6,601通具有两轮或更多轮对话):约895小时的音频和来自5,665个不同主叫号码的328,869轮转录。在整体分类器下,实质性通话主要是掠夺性但合法的潜在客户生成(“垃圾”,约占五分之三),而约七分之一是彻头彻尾的“诈骗”(此快照中为949通)。每通电话都包含逐轮转录、三通道音频、每轮延迟遥测以及多层自动标签,包括由独立人工审查证实的整体诈骗/垃圾/合法判断(二元决策的一致性为75%)。我们描述了收集系统、记录结构以及语料库真实性和标签质量的技术验证,包括代理仅在约5%的参与通话中被识别为非人类。我们还对已建立的诈骗检测方法进行了基准测试,其中在已发布的合成对话上训练的检测器在真实流量上的精确度急剧下降。

英文摘要

Real conversations between fraudsters and their targets are among the most informative artifacts for studying telephone scams, yet also the scarcest: passive honeypots overwhelmingly capture automated messages and hang-ups, large-scale studies characterize call metadata rather than dialogue, and manual scam-baiting does not scale. We present a dataset of real scam-call conversations collected by an active voice-agent honeypot. Dedicated numbers are seeded into the lead-generation channels fraud operations harvest; inbound callers are answered by a low-latency conversational agent that adopts a plausible target persona and sustains the interaction while every call is recorded, transcribed, and automatically labeled. Over an initial 53-day window we captured 10,015 inbound scam and spam calls (6,601 with two or more turns): roughly 895 hours of audio and 328,869 transcribed turns from 5,665 distinct originating numbers. Under a holistic classifier the substantive calls are predominantly predatory-but-legal lead generation ("spam", about three in five), while about one in seven is an outright "scam" (949 in this snapshot). Each call carries a turn-level transcript, three-channel audio, per-turn latency telemetry, and layers of automatic labels, including a holistic scam/spam/legitimate judgment corroborated by independent human review (75% agreement on the binary decision). We describe the collection system, the record structure, and technical validation of the corpus's realism and label quality, including that the agent is recognized as non-human in only about 5% of engaged calls. We also benchmark established scam-detection methods, where detectors trained on published synthetic dialogue collapse in precision on real traffic.

Comments9 pages, 7 figures. Data descriptor. Companion analysis paper forthcoming

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑