发表机构
Hitachi America Ltd.(日立美国有限公司)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究证明SSM安全头的鲁棒认证取决于状态转移矩阵的收缩条件,通过强制收缩将认证比例从41%提升至59%,并实现越狱检测的零样本迁移。
AI 中文摘要
安全头是附加在预训练语言模型上的轻量级分类器,用于在生成之前标记有害输入。其实验检测性能已被研究,但其形式鲁棒性属性在很大程度上仍未探索。我们研究何时可以证明基于状态空间模型(SSM)的安全头在嵌入空间有界扰动内的所有输入上产生相同预测。我们证明答案取决于一个单一条件:状态转移矩阵的$l_\infty$范数必须满足$\norm{A}_\infty<1$(收缩条件),这使得线性时不变分类器能够进行精确的区间边界传播(IBP)认证。当收缩条件成立时,可达输出区间具有有界的稳态宽度,示例可以被认证为鲁棒分类。当条件不成立时,区间随序列长度指数增长,且在任何实际扰动半径下认证都不可能。我们使用铰链惩罚强制收缩,并在有毒评论数据上显示认证比例从41%提高到59%,在$\norm{A}_\infty=1$处出现与理论匹配的尖锐经验相变。将收缩正则化的S4头应用于JailbreakBench上的越狱检测,我们实现了对AdvBench(DR=0.994)和HarmBench(DR=0.988)的零样本迁移。对平均池化的Mamba-130M嵌入进行逻辑回归,在每个检测指标上达到或超过S4头,确认有害意图在嵌入空间中已经是线性可分的。S4安全头的贡献不是优越的判别能力,而是任何基于探针的方法都无法提供的正式认证。
英文摘要
Safety heads are lightweight classifiers attached to pretrained language models for flagging harmful inputs before generation. Their empirical detection performance has been studied, but their formal robustness properties remain largely unexplored. We ask when a State Space Model (SSM)-based safety head can be certified to produce the same prediction for all inputs within a bounded embedding-space perturbation. We prove that the answer turns on a single condition: the $l_\infty$ norm of the state transition matrix must satisfy $\norm{A}_\infty<1$ (the \emph{contraction condition}), which enables exact interval bound propagation (IBP) certification for linear time-invariant classifiers. When the contraction condition holds, the reachable output interval has bounded steady-state width and examples can be certified as robustly classified. When it fails, the interval grows exponentially with sequence length and certification is impossible at any practical perturbation radius. We enforce contraction with a hinge penalty and show on toxic comment data that certified fraction improves from 41\% to 59\%, with a sharp empirical phase transition at $\norm{A}_\infty=1$ matching the theory. Applying a contraction-regularized S4 head to jailbreak detection on JailbreakBench, we achieve a zero-shot transfer to AdvBench (DR=0.994) and HarmBench (DR=0.988). A logistic regression on mean-pooled Mamba-130M embeddings matches or exceeds the S4 head on every detection metric, confirming that harmful intent is already linearly separable in the embedding space. The S4 safety head's contribution is not superior discrimination but the formal certification that no probe-based approach provides.
Comments21 pages, 18 figures, AIMS Workshop @ COLM 2026