发表机构
ShanghaiTech University; Microsoft Research Asia(上海科技大学; 微软亚洲研究院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究针对循环语言模型(LoopLMs)在不同循环深度的安全性问题,提出SafeBridge方法,可降低匹配及跨深度攻击成功率,提升实用性,推动循环计算的安全性对齐。
AI 中文摘要
循环语言模型(LoopLMs)通过在循环步骤中重复使用共享参数,提供了一种参数高效的方法来扩展模型能力。由于每个循环深度都可以独立读出,单个LoopLM在推理深度上暴露了更广泛的输出空间,这引发了一个重要问题:安全性是否在整个循环计算过程中都得到保留。先前的评估表明,更深的循环可以提高对有害查询的安全性,但在越狱攻击下的鲁棒性仍不清楚。因此,我们针对针对不同循环深度的越狱攻击,对LoopLMs进行了全面的安全性评估。我们发现,在更深的推理深度上,攻击成功率会增加,同一查询在不同深度会引发不同的安全行为,针对一个深度构造的攻击可以迁移到其他深度。此外,SFT(监督微调)和偏好对齐无法消除这些安全差距,这促使我们设计一种针对LoopLMs的对齐方法。我们引入了SafeBridge,它结合了轻量级的特定深度共享循环层控制、选择性状态桥接以及跨循环深度的联合安全监督。在不同模型规模、多种攻击方法和安全基准下,SafeBridge大幅降低了匹配深度和跨深度攻击的成功率,同时在保持相当的过度弃权(不执行)行为的情况下,比普通模型提高了通用实用性。我们的结果表明,LoopLM的安全性不能从单个循环深度推断,这促使在循环计算中进行安全性对齐。我们的代码和模型检查点将在论文接收后发布。
英文摘要
Looped Language Models (LoopLMs) provide a parameter efficient approach to scaling model capabilities through repeated use of shared parameters across recurrent steps. Since each recurrent depth can be read out independently, a single LoopLM exposes a broader output space across inference depths, raising an important question: whether safety is preserved throughout recurrent computation. Prior evaluations suggest that deeper recurrence can improve safety on harmful queries, but robustness under jailbreak attacks remains unclear. We therefore conduct a comprehensive safety evaluation of LoopLMs under jailbreak attacks targeting different recurrent depths. We find that attack success can increase at deeper inference depths, the same query can elicit different safety behaviors across depths, and attacks constructed against one depth can transfer to others. Moreover, SFT and preference alignment do not eliminate these safety gaps, motivating an alignment method designed for LoopLMs. We introduce SafeBridge, which combines lightweight depth specific control of shared recurrent layers, selective state bridging, and joint safety supervision across recurrent depths. Across model scales, multiple attack methods, and safety benchmarks, SafeBridge substantially reduces attack success for both matched-depth and cross-depth attacks. It also improves general utility over the vanilla models while maintaining comparable over-refusal behavior. Our results show that the safety of a LoopLM cannot be inferred from a single recurrent depth, motivating safety alignment across recurrent computation. Our code and model checkpoints will be released upon acceptance.
Comments25 pages, 6 figures. Submitted to ICLR 2027