arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.23497cs.AIcs.CL

通过安全方向惩罚缓解推理诱导的不一致性

Mitigating Reasoning-Induced Misalignment via Safety-Direction Penalty

  • University of Toronto(多伦多大学)
  • King Abdullah University of Science and Technology(阿卜杜拉国王科技大学)
  • University of Texas at Austin(德克萨斯大学奥斯汀分校)

机构由 AI 辅助整理,请以论文原文为准。

Yipeng Zhao, Qishun Yang, Shenzhe Zhu, Shu Yang, Di Wang

AI总结:

针对推理微调引发LLM安全退化的RIM问题,本文提出安全方向惩罚(SDP)方法,通过定位安全决策层并惩罚安全方向位移,在Qwen2.5系列模型上实现安全与推理性能的平衡。

AI中文摘要:

推理诱导的不一致性(RIM)是指,在包含数学、代码、带思维链的问题解决等无害内容的推理数据上进行微调时,会导致大型语言模型(LLM)产生有害行为,对LLM推理的安全性构成严重挑战。跨架构、跨规模、跨数据集的检查显示,RIM并非总会出现。以往研究将RIM归因于神经元级纠缠,但未确定该纠缠背后表征空间的几何结构,也未提出训练时的解决方案。本文提供了这两方面的内容:对RIM的表征空间分析,以及安全方向惩罚(SDP),该方法会在推理微调过程中惩罚沿已学习安全方向的移动。分析提取了激活空间中的两个方向:一个编码推理能力,另一个编码安全行为,二者相互耦合:提升推理能力的微调会改变安全表征,且偏移量越大的提示会表现出越严重的安全退化。通过中心核对准(CKA)距离比值和探测法,定位了与该偏移最相关的安全决策层。这些发现为SDP的设计提供了指导:耦合性促使惩罚沿安全方向的位移,而层定位确定了初始范围。当初始范围之外存在补偿性偏移时,相同的诊断方法会指导迭代扩展。在Qwen2.5-3B和7B模型上,SDP在恢复安全性的同时保留了基准推理性能。

英文摘要:

Reasoning-Induced Misalignment, where fine-tuning on reasoning data containing no harmful content, including mathematics, code, and problem-solving with chain-of-thought traces can induce harmful behaviors of LLM, posing a serious challenge to the safety of LLM reasoning. Cross-architecture, cross-scale, and cross-dataset checks show that RIM does not always emerge. Previous work attributed RIM to neuron-level entanglement, but did not identify the geometry of the representation space underlying this entanglement or propose a training-time fix. We provide both: a representation-space analysis of RIM and the Safety-Direction Penalty (SDP), which penalizes movement along a learned safety direction during reasoning fine-tuning. The analysis extracts two activation-space directions, one encoding reasoning ability and the other safety behavior. These directions are coupled: fine-tuning that improves reasoning shifts safety representations, and prompts with larger shifts show larger safety degradation. CKA distance ratios and probes locate the safety-decision layers where this shift is most relevant. These findings guide the design of SDP: the coupling motivates penalizing displacement along the safety direction, and the layer localization sets the initial scope. When the initial scope leaves compensatory shifts beyond the penalized layers, the same diagnostics guide iterative expansion. On Qwen2.5-3B and 7B, SDP restores safety while preserving benchmark reasoning performance.

补充信息

↑