arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.09030cs.AIcs.CLcs.ITcs.LGmath.IT

答案分布轨迹:LLM推理的随机动力学视角

Answer-Distribution Trajectories: A Stochastic-Dynamics View of LLM Reasoning

Mar Gonzàlez I Català, Haitz Sáez de Ocáriz Borde, Davide Murari, Carola-Bibiane Schönlieb, Pietro Liò, George Montañez

首次发表
浏览论文内容

中文总结 AI 辅助

本文提出答案分布轨迹,以随机动力学视角跟踪推理中完整预测分布,区分推理成败机制,为分析评估LLM推理动力学提供丰富框架。

中文摘要 AI 辅助

链式思维推理在模型的输入和最终答案之间提供了结构化的计算过程。然而,它通常通过端点准确性来评估,这忽略了到达该答案所采取的路径。新兴的研究方向利用熵轮廓来解决这一局限性,熵轮廓跟踪推理过程中不确定性的演变,但并未揭示哪些竞争性假设导致了这种不确定性。我们引入了答案分布轨迹,这是一种受随机动力学启发的表示方法,它跟踪模型在推理展开过程中对答案的完整预测分布。作为比端点和熵摘要严格更精细的表示,答案分布轨迹使我们能够通过一个涵盖探索、修正、运动和承诺的动力学推理轮廓来表征轨迹,并区分推理成功和失败的不同动力学机制。在十六个开放权重语言模型和四个推理基准上,我们表明具有相同端点和相似熵轮廓的轨迹可以表现出显著不同的推理动力学。我们进一步发现这些动力学在模型内部和跨模型以及任务之间存在显著差异,不同的目标偏好不同的动力学轮廓。此外,我们表明训练和推理选择会系统地重塑这些轮廓。我们的结果表明,答案分布轨迹为分析和评估LLM推理的动力学提供了一个丰富的框架。

英文摘要

Chain-of-thought reasoning provides a structured computation between a model's input and final answer. Yet it is often evaluated through endpoint accuracy, which ignores the path taken to reach that answer. An emerging line of work addresses this limitation using entropy profiles, which track how uncertainty evolves over the reasoning process but do not reveal which competing hypotheses account for that uncertainty. We introduce answer-distribution trajectories, a stochastic-dynamics-inspired representation that tracks the model's full predictive distribution over answers as reasoning unfolds. As a strictly finer representation than endpoint and entropy summaries, answer-distribution trajectories enable us to characterize a trace through a dynamical reasoning profile spanning exploration, revision, motion, and commitment, and to distinguish different dynamical mechanisms of reasoning success and failure. Across sixteen open-weight language models and four reasoning benchmarks, we show that traces with the same endpoint and similar entropy profiles can exhibit substantially different reasoning dynamics. We further find substantial variation in these dynamics both within and across models and tasks, with different objectives favoring different dynamical profiles. Additionally, we show that training and inference choices systematically reshape these profiles. Our results suggest that answer-distribution trajectories provide a rich framework for analysing and evaluating the dynamics of LLM reasoning.

发表机构

  • University of Cambridge(剑桥大学)
  • Harvey Mudd College(哈维穆德学院)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑