用互信息揭示推理动力学:思维标记是大语言模型推理中的信息峰值
Demystifying Reasoning Dynamics with Mutual Information: Thinking Tokens are Information Peaks in LLM Reasoning
- Gaoling School of Artificial Intelligence, Renmin University of China(中国人民大学北京校区人工智能学院)
- Shanghai Artificial Intelligence Laboratory(上海人工智能实验室)
- University College London, University of London(伦敦大学大学学院)
- Dalian University of Technology(大连理工大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本文从互信息角度揭示大型推理模型中思维标记对应信息峰值并降低预测错误,据此提出两种利用思维标记提升推理性能的简单方法。
AI中文摘要:
大型推理模型(LRMs)在复杂问题求解中展现出令人印象深刻的能力,但其内部推理机制仍未被充分理解。本文从信息论的视角研究LRMs的推理轨迹。通过追踪中间表示与正确答案之间的互信息(MI)在LRM推理过程中的演化,我们观察到一个有趣的MI峰值现象:在LRM的推理过程中,特定生成步骤处的MI会出现突然且显著的上升。我们从理论上分析了这一现象,并证明随着MI增加,模型预测错误的概率会下降。此外,这些MI峰值通常对应于表达反思或转折的标记,例如“Hmm”、“Wait”和“Therefore”,我们将其称为思维标记。随后我们证明,这些思维标记对LRM的推理性能至关重要,而其他标记的影响则微乎其微。基于这些分析,我们提出两种简单而有效的方法,通过精细地利用这些思维标记来提升LRM的推理性能。总体而言,我们的工作为LRM的推理机制提供了新的见解,并给出了提升其推理能力的实用方法。代码可在 https://github.com/ChnQ/MI-Peaks 获取。
英文摘要:
Large reasoning models (LRMs) have demonstrated impressive capabilities in complex problem-solving, yet their internal reasoning mechanisms remain poorly understood. In this paper, we investigate the reasoning trajectories of LRMs from an information-theoretic perspective. By tracking how mutual information (MI) between intermediate representations and the correct answer evolves during LRM reasoning, we observe an interesting MI peaks phenomenon: the MI at specific generative steps exhibits a sudden and significant increase during LRM's reasoning process. We theoretically analyze such phenomenon and show that as MI increases, the probability of model's prediction error decreases. Furthermore, these MI peaks often correspond to tokens expressing reflection or transition, such as ``Hmm'', ``Wait'' and ``Therefore,'' which we term as the thinking tokens. We then demonstrate that these thinking tokens are crucial for LRM's reasoning performance, while other tokens has minimal impacts. Building on these analyses, we propose two simple yet effective methods to improve LRM's reasoning performance, by delicately leveraging these thinking tokens. Overall, our work provides novel insights into the reasoning mechanisms of LRMs and offers practical ways to improve their reasoning capabilities. The code is available at https://github.com/ChnQ/MI-Peaks.