SafeCA:用于文本到视频越狱防御的安全交叉注意力定位与调控
SafeCA: Safe Cross-Attention Localization and Regulation for Text-to-Video Jailbreak Defense
- Nanyang Technological University(南洋理工大学)
- National University of Singapore(新加坡国立大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
SafeCA是一种特征级T2V越狱防御机制,通过分析交叉注意力特征差异提出,可降低主流模型越狱成功率约20%,仅增0.1s推理开销且保持语义一致性,提供架构级可部署防护范式。
AI中文摘要:
文本到视频(Text-to-Video, T2V)生成模型在实际部署中易受越狱攻击,导致生成有害或不当内容。现有防御方法主要依赖输入过滤或重构,不仅会产生高计算延迟,还易扭曲语义。为解决这些问题,本文对干净样本与越狱样本在交叉注意力特征空间的差异进行了实验和系统分析,首次揭示二者在扩散过程中存在累积分离效应及线性可分性逐步增强的趋势。基于该发现,本文提出SafeCA,一种用于安全交叉注意力定位与正则化的特征级防御机制:首先,通过单次推理中从干净提示词收集的交叉注意力特征,经注意力稳定性分析识别关键防御区域与数值;其次,采用带能量归一化的注意力掩码缓解异常激活,并引入轻量级语义空间适配器以重定向异常语义流;此外,通过将特征异常信号反向传播至输入提示词,检测并抑制潜在恶意 token,提升防御在商业模型中的可部署性。实验结果表明,SafeCA在主流T2V模型上将越狱成功率降低约20%,仅增加几乎可忽略的推理开销(+0.1s),且保持良好的文本-视频语义一致性。总体而言,SafeCA为T2V生成模型提供了一种架构级、可部署的防护范式。
英文摘要:
Text-to-Video (T2V) generative models are vulnerable to jailbreak attacks in real-world deployment, leading them to produce harmful or inappropriate content. Existing defense approaches mainly rely on input filtering or reconstruction, which not only incur high computational latency but also tend to distort semantics. To address these issues, we experimentally and systematically analyze the differences between clean and jailbreak samples in the cross-attention feature space, revealing for the first time a cumulative separation effect and a progressively increasing trend of linear separability between the two during the diffusion process. Based on this insight, we propose SafeCA, a feature-level defense mechanism for safe cross-attention localization and regularization. Firstly, we identify key defensive regions and values through attention stability analysis using cross-attention features collected from clean prompts within a single inference. Secondly, SafeCA mitigates anomalous activations via attention masking with energy normalization and introduces a lightweight semantic-space adapter to redirect abnormal semantic flows. Furthermore, we detect and suppress potentially malicious tokens by back-propagating feature anomaly signals to the input cue words, thereby enhancing the deployability of the defense in commercial models. Experimental results show that SafeCA reduces the jailbreak success rate by about 20% on mainstream T2V models, adds almost no inference overhead (+0.1s), and maintains good text-video semantic consistency. Overall, SafeCA provides an architecture-level, deployable protection paradigm for T2V generation models.