arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

思维状态实现内生推理

State of Thought Enables Endogenous Reasoning

Zhiren Gong, Yikun Hou, Zihao Zeng, Ming Xiao, Chau Yuen, Wei Yang Bryan Lim

arXiv 2609.16055首次发表:更新:

发表机构

Nanyang Technological University; KTH Royal Institute of Technology(南洋理工大学; 瑞典皇家理工学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出思维状态(SoT)推理范式,利用模型内部状态驱动推理,通过582参数控制器在冻结骨干上选择性激活历史支持,在多项推理任务中提升准确率并大幅减少令牌与延迟。

AI 中文摘要

测试时计算已成为提升大型语言模型(LLM)能力的主要途径。然而,现有的测试时推理范式严重依赖外部强加的控制,要么通过固定的推理程序,要么通过在受限搜索空间中进行代价高昂的扩展,这限制了两者的泛化能力和效率。我们提出了思维状态(State of Thought, SoT),一种新的推理范式,使LLM能够进行内生推理,由模型的内部推理状态主导推理的展开方式。具体而言,SoT从模型的内部信息传递中提取一个紧凑的动力学-几何状态,并使用一个仅有582个参数的控制器,在冻结的骨干模型上,根据当前推理状态有选择地激活有用的历史推理支持,将推理构建为一种基于证据的状态条件过程,而非外部预设的令牌链。在3个LLM和16个数据集上的定量(1.34倍)、通用(1.62倍)、符号与代码(1.76倍)以及长上下文(2.51倍)推理任务中,SoT持续提升了平均基线准确率,同时将生成的令牌减少了62.6%,端到端延迟降低了44.6%。在2个VLM规模和3个推理任务中,相比推理基线,平均准确率提升了3.8个百分点,相比基于搜索的方法,完成令牌减少了74.9%,延迟降低了73.5%。在受限访问下,SoT在无训练/仅嵌入设置中保留了38.2%/36.5%的平均准确率提升,而仅基于轨迹的评判在3个API模型上达到了84.1%的一致性。总之,内生状态驱动的推理提供了一种可泛化且高效的替代方案。

英文摘要

Test-time compute has emerged as a major approach to improving the capabilities of Large Language Models (LLMs). However, existing test-time reasoning paradigms rely heavily on externally imposed control, either through fixed reasoning programs or through costly expansion in constrained search spaces, limiting both generalization and efficiency. We propose State of Thought (SoT), a new reasoning paradigm that enables endogenous reasoning in LLMs, with the model's internal reasoning state governing how reasoning unfolds. Concretely, SoT extracts a compact dynamics-geometric state from the model's internal information transfer and uses a 582-parameter controller on frozen backbones to selectively activate historical reasoning support useful under the current reasoning state, framing reasoning as a state-conditioned process over evidence rather than an externally prescribed token chain. Across quantitative (1.34x), general (1.62x), symbolic-and-code (1.76x), and long-context (2.51x) reasoning on 3 LLMs and 16 datasets, SoT consistently improves mean-baseline accuracy while reducing generated tokens by 62.6% and end-to-end latency by 44.6%. Across 2 VLM scales and 3 reasoning tasks, it improves mean accuracy by 3.8 points over reasoning baselines, with 74.9% fewer completion tokens and 73.5% lower latency than search-based methods. Under constrained access, SoT retains 38.2%/36.5% mean accuracy gains in training-free/embedding-only settings, while trajectory-only judging reaches 84.1% agreement across 3 API models. Together, endogenous state-driven reasoning provides a generalizable and efficient alternative.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑