arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.17279cs.CRcs.AIcs.MM

在安全边界之间:利用时间一致性破解文本到视频生成模型

Between Safe Boundaries: Exploiting Temporal Consistency for Jailbreaking Text-To-Video Generation Models

  • University of Science and Technology of China(科学技术大学)

机构由 AI 辅助整理,请以论文原文为准。

Xingkai Peng, Jun Jiang, Jiayang Liu, Kejiang Chen, Weiming Zhang

AI总结:

研究针对文本到视频模型的越狱攻击,提出结构化查询高效的BSB框架,利用时间一致性,通过在文本代理空间进行蒙特卡洛树搜索并结合视频级评估来校准结果,实验表明该方法超越现有基线,有效发现模型漏洞。

AI中文摘要:

最近,文本到视频(T2V)模型被广泛部署,人们对其抵御越狱攻击的鲁棒性越来越关注。现有的越狱方法大多改编自文本到图像攻击,应用于T2V系统时有显著缺点。它们未能充分利用视频生成的固有特性——时间一致性,且需要大量视频查询优化,在实际黑盒场景中不可行,对抗性提示搜索也缺乏结构化探索策略。为解决这些限制,我们提出了BSB,一种用于T2V模型的结构化、查询高效的越狱框架。BSB通过将有害意图编码为两个单独无害的边界状态之间的过渡来利用时间一致性。它在更便宜的文本代理空间中进行蒙特卡洛树搜索(MCTS),并定期用稀疏的视频级评估校准搜索结果。我们在包括Veo 3.1、Sora 2、Seedance和Kling v1在内的主流商业T2V模型上进行了全面实验。结果表明,BSB超过了所有现有的越狱基线,在评估模型中,攻击成功率比最强的竞争对手平均相对提高了18.6%。我们的发现确定了时间一致性是T2V模型一个未被充分研究但至关重要的攻击面,并验证了结构化搜索有助于在有限的查询预算下有效发现漏洞。

英文摘要:

Recently, text-to-video (T2V) models have been widely deployed, sparking growing concerns over their robustness against jailbreak attacks. Existing jailbreak methods, mostly adapted from text-to-image attacks, suffer notable drawbacks when applied to T2V systems. They fail to fully leverage temporal consistency, an inherent characteristic of video generation. Besides, these methods demand heavy video query optimization, which is infeasible in practical black-box scenarios. Their adversarial prompt search is also driven by heuristic local signals, lacking principled structured exploration strategies. To tackle these limitations, we propose BSB, a structured, query-efficient jailbreak framework for T2V models. BSB harnesses temporal consistency by encoding harmful intent as the transition between two individually harmless boundary states. Under this paradigm, the attack targets boundary-state pairs whose interpolation tends to produce unsafe intermediate frames during video generation. Directly evaluating all candidate pairs within the video space incurs prohibitive computation cost. Instead, BSB conducts Monte Carlo Tree Search (MCTS) in a cheaper textual proxy space and regularly calibrates search outcomes with sparse video-level evaluations. We conduct comprehensive experiments on mainstream commercial T2V models including Veo 3.1, Sora 2, Seedance and Kling v1. Results show BSB surpasses all existing jailbreak baselines, delivering an average 18.6% relative gain in attack success rate over the strongest competitor across evaluated models. Our findings identify temporal consistency as an understudied yet vital attack surface for T2V models and verify that structured search facilitates effective vulnerability discovery under constrained query budgets.

↑