arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

Multi2AV-Safety:多模态到音视频生成的安全性基准测试

Multi2AV-Safety: Benchmarking Safety in Multimodal-to-Audio-Video Generation

Kaichao Jiang, Changtao Miao, Baiqi Wu, Zhiyuan Lu, Kang Yang, Peiwei Zhao, Junchi Chen, Yunfeng Diao, He Liu, Qi Chu, Tao Gong

arXiv 2608.26535首次发表:更新:

发表机构

University of Science and Technology of China; Zhejiang University; Hefei University of Technology(中国科学技术大学; 浙江大学; 合肥工业大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究推出首个覆盖11种非单例多模态条件配置的音视频生成安全性基准Multi2AV-Safety,发现现有安全守卫存在组合风险感知不足的系统性弱点。

AI 中文摘要

音视频生成正迅速从提示驱动的合成转向多模态条件生成,其中文本、图像、音频和视频可共同塑造生成输出。这一转变改变了安全性评估的性质:有害意图可能不再存在于单一输入中,而是来自跨模态及时序的原本良性或弱有害条件的相互作用。然而,现有的安全性基准大多仍以提示为中心或与固定的条件接口绑定,难以系统地研究这类组合风险。为填补这一空白,我们推出Multi2AV-Safety,据我们所知,这是首个覆盖音视频生成所有11种非单例文本/图像/音频/视频(T/I/A/V)条件配置的安全性基准,包含11024个攻击实例。对Multi2AV-Safety的评估显示,代表性多模态安全守卫在攻击机制和有害证据结构方面存在系统性弱点。我们的评估揭示了两种互补的失效模式:有害语义可由单独良性输入的组合产生;而明确的有害线索在与良性多模态上下文混合时会更难检测。这些结果共同表明,组合风险感知是多模态条件音视频生成安全防护的核心能力缺口:即使所有条件输入都可观测,当前的安全守卫也无法可靠地整合跨模态及时序的安全证据。该数据集将于2026年10月公开发布。

英文摘要

Recent audio-video generators increasingly support joint conditioning on text, images, audio, and video. These capabilities also enable attacks that exploit cross-modal interactions or obscure harmful intent to bypass safeguards and induce harmful audio-video outputs. However, existing generation-safety benchmarks have not kept pace with these advances, providing limited coverage of multimodal input combinations and obscured attack intents. To address these gaps, we introduce Multi2AV-Safety, the first full-coverage red-team benchmark for multimodal-to-audio-video generation, comprising 11,024 attack instances across all 11 non-singleton T/I/A/V conditioning configurations, 4 attack-intent categories, and 5 harm categories. Our evaluation of recent state-of-the-art models, including four multimodal-conditioned audio-video generators and eight safety guards, reveals substantial vulnerabilities in both generation and safeguarding, with multimodal compositional risk and obscured attack-intent risk emerging as two complementary challenges. Guided by these findings, we introduce PerceptGuard, an omni-modal guard integrating compositional-risk and attack-intent supervision through structured risk perception learning. By jointly training rationale generation and safety classification, it learns shared risk representations that enable a safety head to make efficient predictions at inference without rationale decoding, while retaining the ability to generate explanations on demand. Across 34 safety benchmarks, PerceptGuard combines SOTA multimodal safety detection with highly competitive unimodal performance, strengthening input-side safeguards against multimodal attacks on omni models. In particular, it improves safeguarding against the above risks, achieving an overall recall of 86.06\% on Multi2AV-Safety and outperforming GuardReasoner-Omni by 14.56\%.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑