arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

少即是多:通用AIGC音频-视频检测的模态解耦

Less is More: Modality-Decoupling for General AIGC Audio-Video Detection

Jielun Peng, Yabin Wang, Yaqi Li, Jincheng Liu, Xiaopeng Hong, Athanasios V. Vasilakos

arXiv 2607.25543首次发表:更新:

发表机构

Harbin Institute of Technology; University of Agder(哈尔滨工业大学; 阿格德大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对通用AIGC音频-视频检测,现有方法假设不总成立,本文提出DAV-Det系统,通过决策级融合独立建模各模态取证证据,视觉和音频检测器分别利用多粒度表示及门控双分支架构,在挑战赛中排名第一。

AI 中文摘要

生成式人工智能已将视听伪造从以人类为中心的深度伪造迅速扩展到一般场景。现有AIGC检测方法假定视听内容对应,通过发现跨模态不一致来识别伪造。但作者通过实证发现该假设在一般场景中并非始终成立。作者认为决策级融合是比特征级融合更稳健的选择,因此提出DAV-Det,一个解耦的视听AIGC检测系统,从各模态独立建模取证证据。视觉检测器利用多粒度表示捕获空间伪造线索,音频检测器通过门控时间-频谱双分支架构利用时间和频谱不规则性建模声学伪像。该方法在IJCAI-ECAI 2026 DDL 2.0研讨会的通用AIGC音频-视频检测挑战赛中排名第一,最终得分为0.8460。

英文摘要

Generative AI has rapidly expanded audio-visual forgery beyond human-centric deepfakes into general scenes. Existing AIGC detection methods assume audio-visual content correspondence, identifying forgeries by spotting cross-modal inconsistencies. However, we empirically find that this assumption does not consistently hold in general scenarios. We argue that, for general audio-visual AIGC detection, decision-level fusion is a more robust alternative to feature-level fusion. Therefore, we propose DAV-Det, a decoupled audio-visual AIGC detection system that independently models forensic evidence from each modality. The visual detector leverages multi-granularity representations at global, patch, and segment levels to capture spatial forgery cues, while the audio detector exploits both temporal and spectral irregularities via a gated temporal-spectral dual-branch architecture to model acoustic artifacts. Our method ranks 1st in the General AIGC Audio-Video Detection Challenge of the IJCAI-ECAI 2026 DDL 2.0 Workshop, with a final score of 0.8460. Code is available at https://github.com/tuffy-studio/DAV-Det.

CommentsFirst place in the General AIGC Audio-Video Detection Challenge at the IJCAI-ECAI 2026 DDL 2.0 Workshop

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑