大型语言模型中思维链推理的平均场动力学
Mean-Field Dynamics of Chain-of-Thought Reasoning in Large Language Models
浏览论文内容
中文总结 AI 辅助
本研究提出框架,将LLM推理建模为线索图引导式发现过程,用平均场近似推导线索占比的常微分方程,实验验证所得统计规律可复现且能被该方程拟合。
中文摘要 AI 辅助
近年来,具备思维链(Chain-of-Thought)推理能力的大型语言模型(LLMs)已得到广泛应用,对其行为的理论解释有助于加深理解并指导模型优化。本研究引入了一个框架,旨在不简化模型架构也不类比现有物理系统的前提下,探寻LLM推理中的统计规律与理论解释。我们将LLM推理建模为线索图上的引导式发现过程,并利用平均场近似推导出已发现线索占比的一维常微分方程。实验中,线索标记通过学生LLM对教师LLM输出的归一化意外值(normalized surprisal)识别,统计规律则通过对大量思维链推理过程取平均得到。实验表明,所得统计规律在同一数据集内具有可复现性,且可通过求解所提出的理论方程进行拟合。
英文摘要
Large language models (LLMs) with chain-of-thought reasoning have been widely applied in recent years, and theoretical explanations of their behavior may help deepen our understanding and guide model optimization. In this study, we introduce a framework that seeks statistical regularities and theoretical interpretations in LLM reasoning without simplifying the model architecture or making analogies to existing physical systems. We formulate LLM reasoning as a guided discovery process on a clue graph, and derive a one-dimensional ordinary differential equation for the fraction of discovered clues using the mean-field approximation. Experimentally, clue tokens are identified using the normalized surprisal of a student LLM on the outputs of a teacher LLM, and statistical regularities are obtained by averaging over many reasoning chains of thought. Our experiments show that the resulting statistical regularities are reproducible within the same dataset and can be fitted by the solving the proposed theoretical equation.
发表机构
- Tsinghua University(清华大学)
机构由 AI 辅助整理,请以论文原文为准。