arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

更深层推理会损害对齐吗?大型推理模型中对齐崩溃的揭示与缓解

Does Deeper Reasoning Compromise Alignment? Revealing and Mitigating of Alignment Collapse in Large Reasoning Models

Yu-Hang Wu, Yu-Jie Xiong, Henghua Zhang, Bairui Zhang, Jia-Chen Zhang, Shaohua Li

arXiv 2609.08186首次发表:更新:

发表机构

Shanghai University of Engineering Science; The Hong Kong University of Science and Technology; A * STAR(上海工程技术大学; 香港科技大学; 新加坡科技研究局)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文揭示大型推理模型在深度推理下会出现对齐崩溃,提出对齐损失率指标和推理陷阱越狱范式,并设计推理残差对齐防御策略以缓解该问题。

AI 中文摘要

思维链(Chain-of-Thought, CoT)的出现为大型推理模型(Large Reasoning Models, LRMs)奠定了坚实基础。尽管深度推理被广泛认为能增强安全对齐,但在扩展推理下对齐机制的稳定性仍未得到充分探索。本文通过揭示一个关键漏洞来挑战这一普遍观点:深度推理可能引发对齐崩溃(Alignment Collapse)。为严格量化这一现象,我们提出了对齐损失率(Alignment Loss Rate, ALR)指标。我们的实验表明,随着推理深度的增加,ALR显著上升,表明模型对外部扰动的鲁棒性严重下降。利用这一不稳定性,我们提出了一种新颖的越狱范式——推理陷阱(Reasoning Trap, RT)。RT诱导模型进入扩展推理状态,以放大对抗攻击的影响,导致安全能力急剧下降。为阐明这一崩溃背后的机制,我们识别出注意力稀释(Attention Dilution)为根本原因,它源于扩展推理过程与原始输入之间对注意力的竞争。为缓解此问题,我们提出了推理残差对齐(Reasoning Residual Alignment, RRA),一种轻量级防御策略,通过残差连接动态重新强调输入,并与推理过程相集成。

英文摘要

The emergence of Chain-of-Thought (CoT) has established a robust foundation for Large Reasoning Models (LRMs). While deep reasoning is widely believed to enhance safety alignment, the stability of alignment mechanisms under extended reasoning remains underexplored. This paper challenges the prevailing view by revealing a critical vulnerability: Deep Reasoning May Induce Alignment Collapse. To rigorously quantify this phenomenon, we propose the Alignment Loss Rate (ALR) metric. Our experiments demonstrate that as reasoning depth increases, ALR rises significantly, indicating a severe degradation in model robustness against external perturbations. Capitalizing on this instability, a novel jailbreaking paradigm, Reasoning Trap (RT), is proposed. RT induces the model into extended reasoning to amplify the impact of adversarial attacks, leading to a sharp decline in safety capabilities. To elucidate the mechanism behind this collapse, we identify Attention Dilution as the root cause, arising from the competition for attention between the extended reasoning process and the original input. To mitigate this, Reasoning Residual Alignment (RRA), a lightweight defense strategy that dynamically re-emphasizes the input via residual connections integrated with the reasoning process.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑