arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

GenTraceBench:跨预训练与后训练阶段追踪音频深度伪造的基准

GenTraceBench: A Benchmark for Tracing Audio Deepfakes Across Pre- and Post-training Stages

Li Wang, Kunyu Feng, Wan Lin, Dekun Chen, Qinke Ni, Xueyao Zhang, Lei Wang, Jie Shi, Haizhou Li, Zhizheng Wu

arXiv 2609.21738首次发表:更新:

发表机构

The Chinese University of Hong Kong, Shenzhen; Amphion Technology Co., Ltd.; Huawei Technologies Co., Ltd.(香港中文大学(深圳); Amphion科技有限公司; 华为技术有限公司)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

GenTraceBench基准系统评估TTS模型适配后音频深度伪造指纹的稳定性,发现DPO/GRPO保留指纹而部分SFT导致漂移,多镜头注册可显著降低验证错误率。

AI 中文摘要

现代文本转语音(TTS)系统很少以未更改的预训练模型形式部署。它们通常通过监督微调(SFT)或偏好优化(如DPO和GRPO)进行适配。这为音频深度伪造取证提出了一个实际问题:从基础生成器学到的指纹在适配后是否仍然有效?我们提出了GenTraceBench,一个受控基准,涵盖五种TTS架构、16种预训练/后训练变体,以及使用固定文本和说话人提示生成的49,728条话语。在“基于基础训练、基于适配测试”协议下,我们评估了二元检测、闭集归因和开集验证。DPO和GRPO通常保留指纹,而某些SFT和预训练数据变化会导致显著漂移;效应大小在三种取证骨干网络上有所不同。重复训练运行确认了最大的W2V-BERT归因下降,而一个具有可比语音质量的数据混合对照表明,组成变化不一定会导致漂移。在W2V-BERT验证中,多镜头注册将SFT条件下的等错误率(EER)从44.4%降至11.0%,而仅SingNet条件仍保持45%或更高的EER。

英文摘要

Modern text-to-speech (TTS) systems are rarely deployed as unchanged pre-trained models. They are often adapted through supervised fine-tuning (SFT) or preference optimization such as DPO and GRPO. This raises a practical question for audio deepfake forensics: do fingerprints learned from a foundation generator remain valid after adaptation? We present GenTraceBench, a controlled benchmark spanning five TTS architectures, 16 pre-/post-training variants, and 49,728 utterances generated with fixed texts and speaker prompts. Under a train-on-foundation, test-on-adapted protocol, we evaluate binary detection, closed-set attribution, and open-set verification. DPO and GRPO generally preserve fingerprints, whereas some SFT and pre-training-data changes cause substantial drift; effect sizes vary across three forensic backbones. Repeated training runs confirm the largest W2V-BERT attribution drop, while a data-mixture control with comparable speech quality shows that composition change need not cause drift. In W2V-BERT verification, multi-shot enrollment reduces EER for the SFT condition from 44.4% to 11.0%, whereas the SingNet-only condition remains at or above 45% EER.

Comments5 pages, 2 figures, 4 tables. Accepted to the 15th International Symposium on Chinese Spoken Language Processing (ISCSLP 2026)

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑