发表机构
University of Chinese Academy of Sciences; Institute of Computing Technology, Chinese Academy of Sciences; Hunan University; Beijing University of Posts and Telecommunications(中国科学院大学; 中国科学院计算技术研究所; 湖南大学; 北京邮电大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文揭示视频生成中参考图像作为视觉锚点会阻止内容偏离有害意图从而增加安全风险,提出免训练多模态越狱框架DIVA,通过解耦意图和双标准选择实现高攻击成功率,并贡献首个多条件视频生成安全基准TI2VSafetyBench。
AI 中文摘要
视频生成的快速发展已将范式从纯文本驱动转向多条件可控生成,参考图像现被广泛用作条件输入以实现卓越的时空一致性。尽管这些参考图像作为强大的视觉锚点显著增强了可控性,但其对安全性的影响在很大程度上仍未得到探索。在本工作中,我们揭示了视觉锚定效应:通过强制一致性,该机制阻止生成内容偏离原始有害意图,从而消除了模型从有害内容自然逃逸至良性内容的安全出路。因此,视觉锚点固有地增加了安全风险——这就是一致性的代价。基于这一见解,我们提出了通过视觉锚点解耦意图(DIVA),一种针对视频生成的免训练多模态越狱框架,利用了这一漏洞。DIVA将有害意图解耦为静态视觉锚点图像和动态运动文本提示,并采用双标准选择来平衡攻击隐蔽性与语义保留。在多个领先商业平台和主流开源视频生成模型上的大量实验表明,DIVA实现了比现有纯文本方法显著更高的攻击成功率。为促进未来研究,我们还贡献了TI2VSafetyBench,这是首个面向多条件视频生成的安全基准。
英文摘要
The rapid evolution of video generation has shifted the paradigm from pure text-driven to multi-conditional controllable generation, with reference images now widely adopted as conditional inputs to achieve superior spatiotemporal consistency. While these reference images serve as powerful visual anchors that significantly enhance controllability, their impact on safety remains largely unexplored. In this work, we reveal the visual anchoring effect: by enforcing consistency, the mechanism prevents the generated content from drifting away from the original harmful intent, thereby eliminating the model's natural safety escape route from harmful to benign content. Consequently, visual anchors inherently increase the safety risk---this is the price of consistency. Building on this insight, we propose Decoupling Intent via Visual Anchors (DIVA), a training-free multimodal jailbreak framework for video generation that exploits this vulnerability. DIVA decouples harmful intent into a static visual anchor image and a dynamic motion text prompt, and employs dual-criteria selection to balance attack stealthiness with semantic preservation. Extensive experiments across various leading commercial platforms and mainstream open-source video generation models demonstrate that DIVA achieves a substantially higher Attack Success Rate than existing text-only methods. To facilitate future research, we additionally contribute TI2VSafetyBench, the first safety benchmark for multi-conditional video generation.
CommentsThis paper has been accepted to ACM MM 2026