发表机构
University of South Florida(佛罗里达州立大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究针对大语言模型时代TTS和VC系统合成语音与传统基准测试不匹配问题,引入VoxENES 2026基准测试,对八个预训练检测器测试发现性能大幅下降,凸显当前检测器问题,确立该基准为开发音频欺骗对策的实用平台。
AI 中文摘要
现代基于大语言模型的文本转语音(TTS)和语音转换(VC)系统生成的合成语音,与许多传统欺骗基准测试中的生成器不同。这种不匹配造成了时间泛化差距,可能高估真实世界后处理条件下检测器的鲁棒性。我们通过引入VoxENES 2026来弥合这一差距,它是一个包含53628个音频样本的双语(英语和西班牙语)基准测试,使用10种当代语音合成方法生成,并在10种标准化后处理条件下评估。使用VoxENES 2026,我们对八个预训练检测器进行基准测试,发现性能大幅下降:最佳模型总体等效错误率为28.98%,而大多数在现代生成器和干扰下接近或低于随机概率。我们的结果凸显了当前检测器对脆弱特征的依赖,并确立了VoxENES 2026作为开发强大音频欺骗对策的实用测试平台。
英文摘要
Modern LLM-driven text-to-speech (TTS) and voice conversion (VC) systems produce synthetic speech that differs from the generators represented in many legacy spoofing benchmarks. This mismatch creates a temporal generalization gap that can overestimate detector robustness under real-world post-processing conditions. We bridge this gap by introducing VoxENES 2026, a bilingual (English and Spanish) benchmark of 53,628 audio samples generated using 10 contemporary speech synthesis methods and evaluated under 10 standardized post-processing conditions. Using VoxENES 2026, we benchmark eight pretrained detectors without fine-tuning and observe substantial performance degradation: the best model achieves 28.98\% EER overall, while most perform near or below random chance across modern generators and perturbations. Our results highlight the reliance on brittle artifacts in current detectors and establish VoxENES 2026 as a practical testbed for developing robust audio spoofing countermeasures.
CommentsAccepted in InterSpeech 2026