arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

它认为有多难?分析大型语言模型思维链轨迹中感知步骤的推理能量

How Hard Does It Think? Analyzing Step-Aware Reasoning Energy in LLM Chain-of-Thought Trajectories

Hui Wei, Junda Wu, Sheldon Yu, Sizhe Zhou, Yizhu Jiao, Ming Zhong, Bowen Jin, Tong Yu, Shijia Pan, Jiawei Han, Julian McAuley

arXiv 2607.28674首次发表:更新:

发表机构

UC Merced; UC San Diego; UIUC; Adobe Research(加州大学默塞德分校; 加州大学圣迭戈分校; 伊利诺伊大学厄巴纳-香槟分校; 奥多比研究院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究提出SARE框架量化LLM思维链各步骤的推理能量,发现推理能量不均匀、错误轨迹关键节点能量更低,SARE特征性能优于或匹配输出置信度基线,揭示内部几何动态含预测信息。

AI 中文摘要

理解计算资源如何分配到思维链(CoT)推理的各个步骤仍是一个未解决的挑战:现有的可解释性方法依赖输出层面的信号,或将处理深度简化为单一的轨迹级标量,导致步骤级的资源分配情况不透明。我们提出了感知步骤的推理能量(SARE),这是一种几何框架,通过相邻Transformer层的词元隐藏状态的Gram矩阵之间的中心核对齐(CKA),以单个CoT步骤的粒度量化资源分配,捕捉词元间的关系结构,无需特征向量对齐或聚类对应。SARE还通过将CoT轨迹建模为潜在语义状态之间的转换,将这种能量置于推理的语义进展语境中。在六个推理基准和三个开放权重的大型语言模型上,我们发现推理能量在步骤类型之间高度不均匀,表现出轨迹级指标无法察觉的类相变;错误的轨迹在关键推理节点处的能量系统性更低;基于SARE的特征在大多数场景下与基于输出的置信度基线相当或更优,表明内部几何动态编码了超出表面信号的预测信息。

英文摘要

Understanding how computational effort is allocated across individual chain-of-thought (CoT) reasoning steps remains an open challenge: existing interpretability methods rely on output-level signals or collapse processing depth into a single trajectory-level scalar, leaving step-wise effort opaque. We propose Step-Aware Reasoning Energy (SARE), a geometric framework that quantifies effort at the granularity of individual CoT steps via Centered Kernel Alignment (CKA) between Gram matrices of token hidden states across adjacent transformer layers, capturing inter-token relational structure without requiring eigenvector alignment or cluster correspondence. SARE further contextualizes this energy within reasoning's semantic progression by modeling CoT trajectories as transitions among latent semantic states. Across six reasoning benchmarks and three open-weight LLMs, we find that reasoning energy is highly non-uniform across step types, exhibiting phase-like transitions invisible to trajectory-level metrics; incorrect trajectories show systematically lower energy at critical reasoning junctions; and SARE-based features match or outperform output-based confidence baselines in most settings, indicating that internal geometric dynamics encode predictive information beyond surface-level signals.

Comments13 pages, 3 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑