arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.04194cs.CLcs.LG

可解释性≠可理解性:对比思维链推理中的主观判断重要性与实际重要性

Legibility is Not Interpretability: Comparing Judged and Actual Importance in Chain-Of-Thought Reasoning

Kevin Du, Alexander Hoyle, Laura Ruis, Acyr Locatelli

首次发表
浏览论文内容

中文总结 AI 辅助

该研究对比思维链推理的可理解性与实际重要性,发现LLM评判识别高优势步骤的能力有限,微调步骤级批评家仅对错误响应有效,警示勿将推理轨迹可理解性等同于可解释性。

中文摘要 AI 辅助

思维链(Chain-of-Thought,CoT)模型的推理轨迹似乎提供了一个清晰的窗口,可窥见模型得出答案的过程。越来越多的研究将其视为这样的窗口,利用大型语言模型(LLM)评判来诊断错误、评估忠实度,并通过过程奖励模型(process reward models)和生成式批评家提供步骤级监督。这些实践依赖于推理步骤的文本携带其功能角色的信息。但该文本是否实际编码了哪些推理步骤重要的信息?我们将推理步骤的重要性操作化为其优势:包含该步骤时,预期奖励(例如得出正确最终答案)的变化,通过蒙特卡洛回滚(Monte Carlo rollouts)估算。以这些估算值作为基准,我们评估LLM评判能否识别高优势步骤,发现足够强大的LLM可优于流行度基线,但远未达到噪声上限。将模型微调为步骤级批评家对错误响应有显著提升,但对正确响应仍远未达到上限,表明步骤重要性仅能从推理轨迹的文本中部分恢复。我们的发现为日益增多的思维链忠实度研究提供了支撑,该研究警示不要将推理轨迹的可理解性(legibility)等同于可解释性(interpretability),尤其对过程奖励建模具有启示意义。

英文摘要

Reasoning traces from chain-of-thought models appear to offer a legible window into how a model arrives at its answer. A growing body of work treats them as such, using LLM judges to diagnose errors, evaluate faithfulness, and provide step-level supervision via process reward models and generative critics. These practices rely on the text of a reasoning step carrying information about its functional role. But does the text actually encode information about which reasoning steps matter? We operationalize the importance of a reasoning step as its advantage: the change in expected reward, e.g., producing the correct final answer, from including that step, estimated via Monte Carlo rollouts. Basing ground truth on these estimates, we evaluate whether LLM judges can identify high-advantage steps and find that sufficiently capable LLMs can outperform a prevalence baseline but fall well short of a noise ceiling. Fine-tuning a model as a step-level critic yields strong improvement for incorrect responses but remains distant from ceiling for correct responses, suggesting that step importance is only partially recoverable from the text of the reasoning trace. Our findings contribute to a growing body of chain-of-thought faithfulness work that cautions against treating the legibility of reasoning traces as interpretability, especially with implications for process reward modeling.

发表机构

  • ETH Zürich(苏黎世联邦理工学院)
  • MIT(麻省理工学院)
  • Cohere(科here公司)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑