TeleAntiFraud 2.0:一个可刷新、基于画像和音频的电信诈骗检测基准
TeleAntiFraud 2.0: A Refreshable, Profile-Grounded, and Audio-Based Benchmark for Telecom Fraud Detection
另 2 家 · 查看机构详情
- JD Technology(京东科技)
- Chongqing Ant Consumer Finance Co., Ltd., Ant Group(重庆蚂蚁消费金融有限公司,蚂蚁集团)
- Meta
- University of Science and Technology of China(中国科学技术大学)
- Northeastern University(东北大学)
机构由 AI 辅助整理,请以论文原文为准。
浏览论文内容
中文总结 AI 辅助
TeleAntiFraud 2.0提出可刷新、基于画像的音频基准,通过混合树流水线生成近域负例,揭示分类器在近域场景下性能显著下降,确立近域构建和崩溃感知报告为评估核心要求。
中文摘要 AI 辅助
电信诈骗话术演变迅速,且常被设计成类似于日常服务通话,这对基于音频的电信诈骗评估提出了两个关键要求。首先,基准必须纳入新观察到的诈骗模式,同时不覆盖先前建立的测试集。其次,基准必须区分诈骗与合法的近域通话,而非依赖主题分离的负例。我们提出了TeleAntiFraud 2.0,它由我们的混合树反欺诈生成流水线构建,并在月度冻结评估协议下进行评估。该流水线将在线诈骗案例摘要转化为基于画像的场景,通过混合树生成进行扩展,在共享上下文下实现诈骗和非诈骗对话路径,将验证后的对话渲染为角色匹配的语音,并为每个月度评估集冻结生成的音频、标签、提示、清单和来源记录。每个冻结集包含900通中文电话,包括600个诈骗案例和300个近域非诈骗案例。受控文本实验表明,三个分类器在与无关或普通负例评估时达到完美的宏平均F1(Macro-F1),但在近域兄弟负例下降至0.65-0.68。全集音频和自动语音识别加大语言模型(ASR+LLM)评估进一步揭示了类先验捷径、预测崩溃和快照敏感性。这些发现共同确立了近域构建和崩溃感知报告作为在现实混淆条件下评估基于音频的电信诈骗模型的核心要求。随附的研究工件包括构建代码、评估脚本、清单和文档。我们的数据集和代码可在该https URL获取。
英文摘要
Telecom fraud scripts evolve rapidly and are often designed to resemble routine service conversations, creating two key requirements for audio-based telecom-fraud evaluation. First, benchmarks must incorporate newly observed scam patterns without overwriting previously established test sets. Second, they must distinguish fraud from lawful, near-domain calls rather than relying on topic-separated negative examples. We present TeleAntiFraud 2.0, constructed with our Mixed-Tree Anti-Fraud Generation Pipeline and evaluated under a monthly frozen evaluation protocol. The pipeline transforms online fraud-case abstracts into profile-grounded scenarios, expands them through mixed-tree generation, realizes fraud and non-fraud dialogue paths under shared contexts, renders validated dialogues as role-matched speech, and freezes the resulting audio, labels, prompts, manifests, and provenance records for each monthly evaluation set. Each frozen set contains 900 Chinese calls, comprising 600 fraud and 300 near-domain non-fraud cases. Controlled text experiments show that three classifiers achieve perfect macro-averaged F1 (Macro-F1) when evaluated against unrelated or ordinary negatives, but drop to 0.65-0.68 with near-domain sibling negatives. Full-set audio and automatic-speech-recognition plus large-language-model (ASR+LLM) evaluations further reveal class-prior shortcuts, prediction collapse, and snapshot sensitivity. Together, these findings establish near-domain construction and collapse-aware reporting as core requirements for evaluating audio-based telecom-fraud models under realistic confusable conditions. The accompanying research artifact includes the construction code, evaluation scripts, manifests, and documentation. Our dataset and code are available at https://anonymous.4open.science/r/TeleAntiFraud-2_0-EEB2/.