发表机构
School of Cyber Science and Engineering, Huazhong University of Science and Technology; School of Mathematical Sciences, Peking University; School of Software and engineering, Huazhong University of Science and Technology; School of Computer Science, Nanjing University(华中科技大学网络空间科学与工程学院; 北京大学数学科学学院; 华中科技大学软件学院; 南京大学计算机学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出TempJail时序越狱框架,针对图像到视频生成模型的时序漏洞,通过分解恶意提示、受控潜在扰动等方式提升攻击成功率,在Kling等模型上获显著效果。
AI 中文摘要
近年来,图像到视频(I2V)生成模型在主体一致性和时序连贯性方面取得了显著进展,能够生成高质量视频。然而这些进展也带来了新的安全风险。现有研究主要聚焦于涉及单帧违规的越狱攻击,却在很大程度上忽视了视频生成模型特有的时序维度。本文研究了三种攻击场景,揭示了I2V系统中的一种时序漏洞:不安全语义并非来自单帧,而是来自随时间的语义组合。我们进一步确定了此类攻击中的两个关键挑战:时序抽象和语义伪装。为解决这些问题,我们提出了一种用于I2V系统的新型时序越狱框架TempJail。对于时序抽象,我们将目标恶意提示分解为初始帧视觉条件和时序文本指令。对于语义伪装,在图像侧,我们将语义注入建模为扩散采样中的受控潜在扰动,并引入预训练编码器的梯度引导;在文本侧,我们将提示重写为无害的“主体-动作-场景”模板,以绕过安全过滤器,同时保留时序引导。在黑盒推理阶段,这两种模态共同使恶意语义随时间逐步触发。在闭源商业模型(包括Kling、Seedance、Veo和PixVerse)上的实验表明,TempJail在GPT-5.2评估下的攻击成功率较现有最先进方法提升了23.3%,在人工评估下提升了22.0%。我们的代码可在GitHub获取。
英文摘要
In recent years, image-to-video (I2V) generation models have made remarkable progress in subject consistency and temporal coherence, enabling high quality video synthesis. However, these advances also introduce new safety risks. Existing studies mainly focus on jailbreak attacks involving single frame violations, while largely overlooking the temporal dimension unique to video generation models. In this paper, we investigate three attack scenarios and uncover a temporal vulnerability in I2V systems: unsafe semantics may emerge not from a single frame, but from semantic composition over time. We further identify two key challenges in such attacks: temporal abstraction and semantic camouflage. To address these issues, we propose TempJail, a novel temporal jailbreak framework for I2V systems. For temporal abstraction, we decompose a target malicious caption into an initial frame visual condition and a temporal text instruction. For semantic camouflage, on the image side we model semantic injection as controlled latent perturbation in diffusion sampling and introduce gradient guidance from pretrained encoders. On the text side, we rewrite the caption into an innocuous ``subject-action-scene'' template that bypasses safety filters while preserving temporal guidance. In the black-box inference phase, these two modalities jointly enable malicious semantics to be gradually triggered over time. Experiments on closed-source commercial models, including Kling, Seedance, Veo and PixVerse, show that TempJail improves attack success rate over prior state-of-the-art methods by 23.3\% under GPT-5.2 evaluation and 22.0\% under human evaluation. Our codes are available at \href{https://github.com/luqi-glory/TempJail}{GitHub}.
CommentsAccepted by ACM Multimedia 2026 (ACM MM '26)