发表机构
The State Key Laboratory of Blockchain and Data Security, Zhejiang University; Ant Group; Chongqing Ant Consumer Finance Co., Ltd(浙江大学区块链与数据安全全国重点实验室; 蚂蚁集团; 重庆蚂蚁消费金融有限公司)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对大型推理模型安全对齐中的虚假捷径问题,提出DeShortcut-Align框架,通过敏感性归因、对比增强和反事实正则化解耦格式与词汇捷径,显著提升鲁棒性并减少过度拒答。
AI 中文摘要
通过监督微调(SFT)和强化学习(RL)对大型推理模型(LRMs)进行安全对齐,往往能获得近乎完美的安全分数,然而这种表面上的成功却以严重的过度拒答和通用能力下降为代价。通过系统的实证分析,我们发现这些失败与虚假捷径的学习密切相关,而非稳健的意图敏感安全评估。具体而言,我们识别出两种主要的捷径:格式捷径,即拒答行为过度绑定于安全对齐语料中频繁出现的结构提示模板;以及词汇捷径,即敏感关键词在良性查询上反射性地触发拒答。为减轻对这些捷径的依赖,我们提出了DeShortcut-Align,一种捷径解耦对齐框架,旨在减少对表面线索的依赖。DeShortcut-Align在三个协调阶段运作:(1)拒答敏感性归因,通过掩蔽输入标记来量化其对最终拒答响应分布的影响;(2)归因引导的对比增强,利用高敏感性标记构建良性对比样本以缓解词汇捷径;(3)反事实一致性正则化,通过注意力遮蔽构建模板消融状态,以在SFT和RL中强制决策一致性,减轻格式捷径依赖。在7B和14B模型上的实验表明,DeShortcut-Align显著提高了对模板剥离绕过攻击的鲁棒性(性能下降减少高达72%),大幅降低了超过58%的过度拒答,并更好地保留了通用推理能力,从而减轻了安全训练中常见的对齐税。
英文摘要
Safety alignment of large reasoning models (LRMs) via supervised fine-tuning (SFT) and reinforcement learning (RL) often yields near-perfect safety scores, yet this apparent success comes at the cost of severe over-refusal and degraded general capabilities. Through systematic empirical analysis, we find that these failures are closely associated with the learning of spurious shortcuts rather than robust intent-sensitive safety evaluation. Specifically, we identify two dominant shortcuts: formatting shortcuts, where refusal behaviors are overly bound to structural prompt templates that frequently appear in safety alignment corpora; and lexical shortcuts, where sensitive keywords reflexively trigger refusals on benign queries. To mitigate reliance on these shortcuts, we propose DeShortcut-Align, a shortcut-decoupling alignment framework that reduces dependence on superficial cues. DeShortcut-Align operates across three coordinated stages: (1) Refusal Sensitivity Attribution, which masks input tokens to quantify their impact on the final refusal response distribution; (2) Attribution-Guided Contrastive Augmentation, which constructs benign contrastive samples using high-sensitivity tokens to mitigate lexical shortcuts; and (3) Counterfactual Consistency Regularization, which constructs template-ablated states via attention blinding to enforce decision consistency across SFT and RL, mitigating formatting shortcut dependence. Experiments on 7B and 14B models demonstrate that DeShortcut-Align significantly improves robustness against template-stripping bypass attacks (reducing performance drops by up to 72%), substantially reduces over-refusal by over 58%, and better preserves general-purpose reasoning capabilities, thereby mitigating the alignment tax commonly observed in safety training.
Comments35 pages, 7 figures