arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.24639eess.ASeess.SP

探究语音的浊音与清音区域用于音频深度伪造检测

Investigating voiced and unvoiced regions of speech for audio deepfake detection

Ganesh Sivaraman, Hemlata Tak, Elie Khoury

首次发表
浏览论文内容

中文总结 AI 辅助

本研究探究语音浊音与清音区域在音频深度伪造检测中的作用,基于图注意力的AASIST系统分别训练后发现,清音区域检测效果更优,融合浊音区域后性能进一步提升,在MLAAD数据集上取得5.82%的等错误率。

中文摘要 AI 辅助

基于深度神经网络的深度伪造检测系统在基准数据集和竞赛中已达到很高的准确率,但大多数模型缺乏可解释性,难以从网络中提取能说服人类评估者信任其决策的推理依据。人类常依赖非自然的音高抖动、机械语调、声学伪影、发音不自然的摩擦音等声学线索来判断合成音频的质量。本研究探究语音的浊音与清音区域在区分合成语音与真实语音中的作用,采用信号周期性度量将语音分析为浊音与清音分量,再基于图注意力的AASIST检测系统分别在各分量上独立训练。该研究在MLAAD数据集上对比使用浊音与清音分量的深度伪造检测系统准确率并分析结果,结果显示清音区域在区分合成(深度伪造)语音与真实语音方面尤为有效,等错误率为6.62%;通过分数级融合结合语音区域后,整体性能进一步提升,等错误率达5.82%,较使用完整音频的基线系统相对提升49%。

英文摘要

Deep neural network based deepfake detection systems have achieved high levels of accuracy on benchmark datasets and competitions. However, most models lack interpretability. It is challenging to extract reasoning from the network that can convince the human evaluator to trust the decision. Humans often rely on acoustic cues like unnatural pitch jitter, robotic intonation, acoustic artifacts, and unnatural sounding fricatives to judge the quality of the synthetic audio. This study explores the role played by the voiced and unvoiced regions of speech in discriminating synthetic from bonafide speech. A measure of signal periodicity is used to analyze speech into voiced and unvoiced components. Then, the graph attention based AASIST detection system is trained independently on each component. This work compares the accuracy of deepfake detection system using voiced and unvoiced components and analyzes the results on the MLAAD dataset. Our results show that unvoiced regions are particularly more effective in distinguishing synthetic (deepfake) speech from bonafide, and achieves an equal error rate of 6.62%. When combined with voice regions through score-level fusion, the overall performance improves further, yielding a 5.82% EER, a relative improvement of 49% over the baseline system that uses the full audio.

补充信息

↑