发表机构
Ewha W. University; KAIST AI(梨花女子大学; 韩国科学技术院人工智能学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究通过F5-TTS-BigVGAN流水线追踪检测证据来源,发现声学模型更新可减少固定检测器的检测证据,而检测器适应能恢复检测能力,显著降低EER。
AI 中文摘要
最近的音频深度伪造检测器能够区分真实语音与合成语音,但文本到语音系统的哪个阶段提供了检测证据仍不清楚。我们通过F5-TTS-BigVGAN流水线中的受控重合成和检测器适应来解决这一问题。由于声码器对真实梅尔谱的重建本身可能与源话语可分离,我们固定声码器,并将检测器分离度的较大变化追溯到声学生成。对声学模型进行对抗性微调,其目标中不包含检测器,在质量相当的情况下,提高了针对固定检测器的等错误率(EER)。然而,仅对微调模型的VCTK输出进行检测器适应,将其LibriSpeech EER从19.42%降至7.46%,并改善了对未见的基础F5-TTS输出的检测。这些结果表明,声学模型的更新可以减少固定检测器可用的检测证据,而检测器适应则在此流水线中保持更新后的输出可检测。
英文摘要
Recent audio deepfake detectors separate bona fide speech from synthetic speech, yet it remains unclear which stage of a text-to-speech system supplies the detection evidence. We address this with controlled resynthesis and detector adaptation in an F5-TTS-BigVGAN pipeline. Since vocoder reconstruction of a real mel can itself be separable from the source utterance, we fix the vocoder and trace the larger change in detector separation to acoustic generation. Adversarially fine-tuning the acoustic model, with no detector in its objective, raises EER against fixed detectors at comparable quality. However, adapting a detector only on the tuned model's VCTK outputs lowers its LibriSpeech EER from 19.42% to 7.46% and improves detection of unseen base F5-TTS outputs. These results suggest that acoustic-model updates can reduce the detection evidence available to fixed detectors, while detector adaptation keeps the updated outputs detectable in this pipeline.
CommentsSubmitted to ICASSP 2027