首个令牌至关重要:理解大型推理模型中的安全崩溃
First Token Matters: Understanding Safety Collapse in Large Reasoning Models
- Harbin Institute of Technology(哈尔滨工业大学)
- Guangzhou University(广州大学)
- City University of Macau(澳门城市大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本文发现大型推理模型在有害查询下首个生成令牌处出现拒绝信号崩溃(ORC),提出轻量级推理时干预SafeToken,通过注入安全锚点缓解该问题,提升安全性并保留推理能力。
AI中文摘要:
大型推理模型(LRMs)展现出强大的问题解决能力,然而在处理有害查询时,其安全对齐往往退化。现有的提升安全性的方法主要依赖于额外的训练或偏好优化,而对安全失败背后的内部机制理解有限。在本工作中,我们通过令牌级位置分析来研究拒绝动态中的这种失败,并识别出推理起始处的一个局部脆弱点,我们将其称为“起始拒绝崩溃”(ORC)。我们发现,在有害查询下,LRMs的拒绝相关信号在第一个生成的令牌处急剧下降,这与不安全响应的生成相关。受此发现启发,我们提出了SafeToken,一种轻量级的推理时干预方法,它在推理起始处精确注入一个学习得到的连续安全锚点。尽管仅更新单个令牌嵌入,SafeToken能有效缓解ORC,在有害查询基准上提升安全性,并在很大程度上保留推理效用。这些结果表明,LRMs中的安全失败可能源于从理解到生成的关键过渡中的瞬时崩溃。
英文摘要:
Large Reasoning Models (LRMs) exhibit strong problem-solving abilities, yet their safety alignment often degrades when handling harmful queries. Existing approaches to improving safety largely rely on additional training or preference optimization, while offering limited understanding of the internal mechanisms behind safety failures. In this work, we investigate this failure through a token-level positional analysis of refusal dynamics and identify a localized vulnerability at the onset of reasoning, which we term Onset Refusal Collapse (ORC). We find that the refusal-related signal of LRMs drops sharply at the first generated token under harmful queries, which is associated with unsafe response generation. Motivated by this finding, we propose SafeToken, a lightweight inference-time intervention that injects a learned continuous safety anchor precisely at reasoning onset. Despite updating only a single token embedding, SafeToken effectively mitigates ORC, improves safety on harmful-query benchmarks, and largely preserves reasoning utility. These results suggest that safety failures in LRMs can arise from a transient breakdown at the critical transition from understanding to generation.