发表机构
Electronics and Telecommunications Research Institute; University of Science and Technology(电子通信研究院; 科技大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究比较域不变韵律特征与自监督模型(HuBERT、wav2vec2.0)在跨域语音钓鱼检测中的表现,发现域不变特征支持零样本部署,而自监督方法性能更高但需真实样本与计算资源。
AI 中文摘要
语音钓鱼检测面临三个关键挑战:由于隐私限制,真实的犯罪录音不可用;即便可用,也仅有少量样本,不足以进行微调;此外,需要轻量级的仅基于声学的检测方法,作为大型自监督模型的替代方案。我们通过跨域评估(在基于场景的演员录音上训练,在真实犯罪电话上测试)比较了域不变韵律特征与自监督表示(HuBERT、wav2vec2.0)。域不变韵律特征在零样本下达到69.5%的F1分数,在5样本学习下达到71.0%。HuBERT取得了最高性能(5样本下F1为94.2%),而wav2vec2.0表现出以精确度为导向的检测特征(5样本下F1为90.2%,精确度为99.4%)。这些发现揭示了基本的权衡:域不变特征在无真实数据时支持零样本部署,而自监督方法虽能达到更高性能,但需要真实样本和计算资源。
英文摘要
Voice phishing detection faces three critical challenges: real criminal recordings are unavailable due to privacy constraints; when available, only a handful of samples exist, insufficient for fine-tuning; and lightweight acoustic-only detection is needed as an alternative to large self-supervised models. We compare domain-invariant prosodic features and self-supervised representations (HuBERT, wav2vec2.0) through cross-domain evaluation-training on scenario-based actor recordings and testing on authentic criminal calls. Domain-invariant prosodic features achieve 69.5% F1 zero-shot and 71.0% with 5-shot learning. HuBERT achieves highest performance (94.2% F1, 5-shot), while wav2vec2.0 exhibits a precision-oriented detection profile (90.2% F1 with 99.4% precision, 5-shot). These findings reveal fundamental trade-offs: domain-invariant features enable zero-shot deployment when no real data exists, while SSL methods achieve higher performance but require real samples and compute.
Comments5 pages, 2 figures. Accepted at INTERSPEECH 2026