arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

Teffic-Audio:区分事实与虚构

Teffic-Audio: Tell Fact from Fiction

Wan Lin, Li Wang, Jindong Wang, Kunyu Feng, Zhizheng Wu

arXiv 2607.28351首次发表:更新:

发表机构

Amphion Team(Amphion团队)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

Teffic-Audio是一款基于Conformer架构的通用语音深度伪造检测系统,通过特定训练方案提升泛化能力,在Speech-DF-Arena测试集上的表现优于现有公开系统,为该领域提供实用参考。

AI 中文摘要

语音深度伪造检测的范围已随越来越多样的伪造机制而扩大,这些机制包括语音合成、语音转换、声码器重建以及神经编解码器重合成。生成的伪造痕迹还会因源语音、录音环境和传输信道的差异而进一步改变。这种变异性使得在异质条件下实现鲁棒泛化成为实用检测系统的核心要求。本报告介绍了Teffic-Audio,这是一款专为综合评估环境设计的通用语音深度伪造检测系统。Teffic-Audio采用简单的检测器架构,由基于Conformer的语音编码器、多头注意力统计池化和二分类器组成。该系统不依赖额外的架构复杂性,而是通过其训练方案提升泛化能力,该方案整合了多源数据、针对攻击和源的平衡采样以及多样化音频增强。仅使用开源数据进行训练的Teffic-Audio在Speech-DF-Arena的14个测试集上实现了1.454%的等错误率(EER),在排行榜上优于所有当前公开系统,还在5个单独测试集上获得了最低EER,且与更大的领先系统相比,展现出良好的性能-复杂度权衡。总体而言,Teffic-Audio为通用语音深度伪造检测提供了强大且实用的参考系统。

英文摘要

Speech deepfake detection has expanded in scope with increasingly heterogeneous spoofing mechanisms, including speech synthesis, voice conversion, vocoder reconstruction, and neural-codec resynthesis. The resulting spoofing artifacts can be further shaped by variability in source speech, recording environments, and transmission channels. This variability makes robust generalization across heterogeneous conditions a central requirement for practical detection systems. This report presents Teffic-Audio, a general speech deepfake detection system designed for comprehensive evaluation environment. Teffic-Audio adopts a straightforward detector architecture consisting of a Conformer-based speech encoder, multi-head attentive statistics pooling, and a binary classifier. Rather than relying on additional architectural complexity, the system improves generalization through its training recipe, which integrates multi-source data, attack- and source-balanced sampling, and diverse audio augmentation. Trained only with open-source data, Teffic-Audio achieves a pooled EER of 1.454% on the 14 test sets of Speech-DF-Arena, outperforming all currently public systems on the leaderboard. It also obtains the lowest EER on five individual test sets and shows a favorable performance-complexity trade-off compared with larger leading systems. Overall, Teffic-Audio provides a strong and practical reference system for general speech deepfake detection.

Comments16 pages, 1 figure, 7 tables. Technical report. Project page: https://tefficlabs.com/teffic-audio

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑