arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

大型语言模型推理的解读:基于部分信息分解

Interpreting Reasoning of Large Language Models via Partial Information Decomposition

Barproda Halder, Qiuyi Zhang, Sanghamitra Dutta

arXiv 2610.00571首次发表:更新:

发表机构

University of Maryland, College Park; Elorian AI(马里兰大学帕克分校; 埃洛里安人工智能公司)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

提出SLIDER框架,利用部分信息分解评估大型推理模型的推理质量,通过Step-RRI和Trajectory-RRI检测冗余,提升效率并用于数据选择微调。

AI 中文摘要

大型推理模型(LRMs)在解决复杂数学问题方面取得了显著进步,但常常产生冗长、重复或错误的推理轨迹。在本工作中,我们引入了一个新的可解释性框架SLIDER,用于评估推理过程的质量。SLIDER利用信息论中一个新兴的研究方向——部分信息分解,将两个连续推理步骤之间关于最终答案的信息分解为非负分量:独特信息(来自先前步骤或当前步骤)、冗余信息和协同信息。基于这一分解,我们提出了*逐步重复推理指数(Step-RRI)*,这是一个有理论基础的度量,用于评估当前步骤$S_i$中与答案相关的信息是否主要与过去步骤$S_{<i}$冗余,相对于其独特和协同贡献。为了评估Step-RRI在检测重复性方面的有效性,我们将SLIDER应用于PRMBench数据集的冗余类别,其中Step-RRI在逐步冗余识别准确率上比嵌入相似性和信息增益基线提高了超过10个百分点。接下来,我们定义了*轨迹重复性指数(Trajectory-RRI)*,这是单个推理轨迹重复性的聚合度量。为了展示其实际相关性,我们表明在QwQ-32B、DeepSeek-R1-Distill-Qwen-32B和GPT-4.1上,平均Trajectory-RRI与实际推理长度强相关,这激励了其作为提高推理效率的信号的使用。最后,我们引入了*基于Trajectory-RRI的数据选择用于微调*,证明基于Trajectory-RRI选择训练数据可以提高微调模型的推理效率,同时基本保持其任务性能。

英文摘要

Large reasoning models (LRMs) have achieved substantial improvements in solving complex mathematical problems, but often produce lengthy, repetitive, or erroneous reasoning trajectories. In this work, we introduce a new interpretability framework, SLIDER, to evaluate the quality of the reasoning process. SLIDER leverages an emerging body of work from information theory called Partial Information Decomposition to disentangle the information about the final answer between two consecutive reasoning steps into non-negative components: unique information (in preceding steps or current step), redundant information, and synergistic information. Building on this decomposition, we propose the *Step-wise Repetitive Reasoning Index (Step-RRI)*, a theoretically grounded measure that assesses whether the answer-relevant information in the current step $S_i$ is predominantly redundant with the past steps $S_{<i}$, relative to its unique and synergistic contributions. To evaluate the effectiveness of Step-RRI in detecting repetitiveness, we apply SLIDER to the redundancy class of the PRMBench dataset where Step-RRI improves step-level redundancy identification accuracy by over $10$ points compared to embedding-similarity and information-gain baselines. Next, we define *Trajectory-RRI*, an aggregate measure of repetitiveness for an individual reasoning trajectory. To demonstrate its practical relevance, we show that average Trajectory-RRI strongly correlates with actual reasoning length across QwQ-32B, DeepSeek-R1-Distill-Qwen-32B, and GPT-4.1, motivating its use as a signal for improving reasoning efficiency. Finally, we introduce *Trajectory-RRI-guided data selection for fine-tuning*, demonstrating that selecting training data based on Trajectory-RRI can improve a fine-tuned model's reasoning efficiency while largely preserving its task performance.

CommentsAccepted at ICLR 2026 Workshop on Logical Reasoning of Large Language Models

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑