发表机构
Mohamed bin Zayed University of Artificial Intelligence; National University of Singapore; Ho Chi Minh City University of Technology (HCMUT), VNU-HCM(穆罕默德·本·扎耶德人工智能大学; 新加坡国立大学; 胡志明市理工大学(越南国立大学胡志明市分校))
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出StreamFraudNet,利用冻结的自监督语音编码器和循环时间建模,从原始电话音频中增量检测诈骗,在受控基准上达到0.9953的ROC-AUC,且无需转录或时间标注。
AI 中文摘要
我们研究从原始电话音频中进行弱监督的增量式电信诈骗检测,其中训练仅提供对话级别的标签,且预测必须在通话结束前更新。我们引入了StreamFraudNet,它通过重叠的有界上下文窗口处理传入的音频,使用冻结的自监督语音编码器、循环时间建模以及学习到的潜在窗口分数聚合。在一个受控的英语基准上,StreamFraudNet达到了0.9953的ROC-AUC,显著优于声学和均值池化基线,同时与强大的全局时间模型保持竞争力。该模型在10秒音频后产生首次预测,每2秒更新一次,并且在评估所用的服务器硬件上运行速度快于实时。消融实验表明,循环时间上下文是性能的主要贡献因素。这些结果表明,诈骗风险可以从原始语音中增量评分,而无需转录或时间标注,同时强调了需要延迟感知的训练以改进早期预测。
英文摘要
We study weakly supervised incremental telecom fraud detection from raw telephone audio, where training provides only conversation-level labels and predictions must be updated before a call ends. We introduce StreamFraudNet, which processes incoming audio through overlapping bounded-context windows using a frozen self-supervised speech encoder, recurrent temporal modeling, and learned aggregation of latent window scores. On a controlled English benchmark, StreamFraudNet achieves a ROC--AUC of \(0.9953\), significantly outperforming acoustic and mean-pooling baselines while remaining competitive with strong global temporal models. The model produces its first prediction after 10 seconds of audio, updates every 2 seconds, and operates faster than real time on the evaluated server hardware. Ablations identify recurrent temporal context as the principal contributor to performance. These results demonstrate that fraud risk can be scored incrementally from raw speech without transcripts or temporal annotations, while highlighting the need for latency-aware training to improve early prediction.