当推理偏离正轨:不受控推理的注意力动态
When Reasoning Goes Astray: Attention Dynamics of Uncontrolled Reasoning
浏览论文内容
中文总结 AI 辅助
本文提出RADAR方法,通过动态注意力响应实时分析大型推理模型的推理状态,识别并修正异常注意力分布,以抑制不受控推理导致的冗余循环,同时保留良性性能。
中文摘要 AI 辅助
大型推理模型(LRMs)通过扩展推理提升了复杂任务上的性能,然而同一过程也可能退化为冗余验证和持续生成循环。这种不受控的推理增加了推理成本,并带来资源耗尽和服务降级的风险。然而,现有的缓解措施大多截断长输出或对表面重复做出反应,因而无法区分正常思考与不受控推理,也无法解释良性推理如何退化为有害行为。在本文中,我们将LRM生成操作化为四种状态,并进一步引入基于动态注意力响应的推理状态分析(RADAR),该方法能够实时识别当前推理状态,并刻画有效反思如何发展为不受控生成。在RADAR分析的指导下,我们进一步将异常的注意力分布重新对齐到正常请求中观察到的模式,并考察这种修正对过度反思和持续循环的影响。时间分析表明,不受控推理的特征是注意力分布偏离正常生成,且异常趋势在重复开始之前就已可检测。通过注意力重对齐修正这些偏差,能够持续减少循环,同时基本保留良性性能。综上,RADAR提供了关于推理如何变得不受控的机制性解释,为识别关键失败阶段和设计针对性运行时干预提供了可操作的指导。
英文摘要
Large reasoning models (LRMs) improve performance on complex tasks through extended reasoning, yet the same process can degenerate into redundant verification and persistent generation loops. Such uncontrolled reasoning increases inference cost and creates risks of resource exhaustion and service degradation. However, existing mitigations largely truncate long outputs or react to surface repetition, and thus fail to distinguish normal thinking from uncontrolled reasoning or explain how benign reasoning degenerates into harmful behavior. In this paper, we operationalize LRM generation as four states and further introduce Reasoning-state Analysis via Dynamic Attention Responses (RADAR), which identifies the current reasoning state in real time and characterizes how effective reflection can develop into uncontrolled generation. Guided by RADAR's analysis, we further realign abnormal attention distributions toward patterns observed in normal requests and examine how this correction affects excessive reflection and persistent looping. Temporal analyses show that uncontrolled reasoning is characterized by attention distributions that deviate from normal generation, with abnormal trends becoming detectable before repetition begins. Correcting these deviations through Attention Realignment consistently reduces looping while largely preserving benign performance. Together, RADAR provide a mechanistic account of how reasoning becomes uncontrolled, offering actionable guidance for identifying critical failure stages and designing targeted runtime interventions.
发表机构
- Beijing University of Posts and Telecommunications(北京邮电大学)
- Wuhan University(武汉大学)
- JIUTIAN Research(中移九天)
- Nanyang Technological University(南洋理工大学)
- Chongqing University of Posts and Telecommunications(重庆邮电大学)
机构由 AI 辅助整理,请以论文原文为准。