arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.00493cs.SD

用于全类型音频深度伪造检测的隐藏域路由

Hidden-Domain Routing for All-Type Audio Deepfake Detection

  • OPPO AI Center(OPPO人工智能中心)

机构由 AI 辅助整理,请以论文原文为准。

Yifan Gao, Yao Tian, Hongbin Suo, Haonan Lu

AI总结:

针对全类型音频深度伪造检测推理时无音频类型的问题,提出闭域路由系统,先恢复隐藏音频域再分支解释分数,在AT-ADD Track2任务中获96.10%宏F1值且排名第一。

AI中文摘要:

全类型音频深度伪造检测需要对语音、环境声、歌声和音乐做出真实性判断,但推理时无法获取音频类型。在AT-ADD Track2任务中,该设置形成了隐藏音频域条件:真实/伪造的二分类标签在各域间一致,但表征结构与检测器分数行为随音频类型变化。本文提出一种闭域路由系统,先恢复隐藏的音频域,再在选定分支内解释检测器分数。AudioType-BEATs-6s路由器从6秒音频窗口估计音频类型;语音输入由Speech-XLSR专家处理,而环境声、歌声和音乐则依赖基于EAT的通用音频专家,采用分支本地分数解释。通过开发集表征分析、路由器家族对比及组件结果,验证了音频域分离效果及不同音频类型检测器的互补优势。在官方AT-ADD Track2最终评估中,该系统取得96.10%的Track2宏F1值,在最终排行榜排名第一,语音、环境声、歌声、音乐的类型宏F1值分别为88.07%、98.18%、99.07%、99.08%。这些结果支持在全类型音频深度伪造检测中,先恢复隐藏音频域再解释检测器分数的思路。

英文摘要:

All-type audio deepfake detection requires authenticity decisions across speech, environmental sound, singing voice, and music, while the audio type is unavailable at inference time. In AT-ADD Track2, this setting creates a hidden audio-domain condition: the binary real/fake label is shared across domains, but representation structure and detector-score behavior vary with audio type. We present a closed-condition routed system that first recovers the hidden audio domain and then interprets detector scores within the selected branch. The AudioType-BEATs-6s Router estimates audio type from a 6-second window; speech inputs are handled by the Speech-XLSR Expert, while sound, singing, and music rely on EAT-based general-audio experts with branch-local score interpretation. Development-set representation analysis, router-family comparisons, and component results show audio-domain separation and complementary detector strengths across audio types. On the official AT-ADD Track2 final evaluation, the system achieves 96.10% Track2 Macro-F1 and ranks first on the final leaderboard, with type-wise Macro-F1 scores of 88.07%, 98.18%, 99.07%, and 99.08% for speech, sound, singing, and music, respectively. These results support recovering the hidden audio domain before interpreting detector scores in all-type audio deepfake detection.

补充信息

↑