arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

追踪语音合成器欺骗人类的趋势

Tracking the Trend in How Speech Synthesizers Deceive People

Milan Šalko, Anton Firc, Kamil Malinka, Vojtěch Staněk, Martin Perešini, Filip Pleško, Jakub Reš

arXiv 2608.19959首次发表:更新:

发表机构

Brno University of Technology(布尔诺理工大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究对比了三款不同年份的现代语音合成器的欺骗性,发现人类和检测器对部分伪造语音的检测准确率极低,仅靠人类感知不可靠,需推进相关验证与检测技术。

AI 中文摘要

语音合成技术的进步使得深度伪造音频的逼真度大幅提升。早期研究报告的人类检测准确率为70%-80%,但这些研究主要依赖较旧的合成器。我们将2019年、2022年和2024年发布的三款选定语音合成工具,与82名IT专业人员的人类检测表现进行对比,并在相同的测试材料上,将人类的表现与六个预训练检测器的性能进行基准测试。对于完全合成语音(完整伪造内容),尽管听众被明确告知存在深度伪造内容,F1分数仍从RTVC和YourTTS的约90%降至ElevenLabs的48%。对于部分伪造(仅改变话语中的一句话),严格准确率降至9%,且听众将合成句子归类为真实句子的比例达77%。人类和检测器的失败方式互补,且两者均无法可靠定位短时间的操纵。此外,听众越来越多地将真实语音误标为伪造,削弱了对未被操纵音频的信任。这些发现表明,在所选的现代合成器和部分伪造场景下,仅靠人类感知不可靠,推动了程序验证、来源溯源、水印技术及片段级检测的发展。

英文摘要

Advances in speech synthesis have made deepfake audio highly realistic. Earlier studies reported 70-80% human detection accuracy, but relied primarily on older synthesizers. We compare human detection for three selected voice synthesis tools released in 2019, 2022, and 2024 with 82 IT professionals, and benchmark humans against six pretrained detectors on the same material. For fully synthetic speech (full spoofs), the F1 score drops from about 90% for RTVC and YourTTS to 48% for ElevenLabs, although listeners were explicitly warned that deepfakes were present. For partial spoofing, where only one sentence of an utterance is altered, strict accuracy falls to 9%, and listeners classify the synthetic sentence as bona fide 77% of the time. Humans and detectors fail in complementary ways, and neither reliably localizes short manipulations. Additionally, listeners increasingly mislabel bona fide speech as fake, eroding trust in unmanipulated audio. These findings show that human perception alone is unreliable for the selected modern and partial-spoof conditions and motivate procedural verification, provenance, watermarking, and segment-level detection.

CommentsAccepted at the 6th Symposium on Security and Privacy in Speech Communication (SPSC 2026)

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑