arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2602.08332cs.CLcs.AI

基于监督思考状态的潜在推理

Latent Reasoning with Supervised Thinking States

  • Google Research(谷歌研究)

机构由 AI 辅助整理,请以论文原文为准。

Ido Amos, Avi Caciularu, Mor Geva, Amir Globerson, Jonathan Herzig, Lior Shani, Idan Szpektor

更新

AI总结:

本文提出Thinking States方法,通过在输入处理过程中生成思考token,实现更高效的潜在推理,实验表明其在多个推理任务中表现优异,且在延迟和长序列外推方面优于CoT。

AI中文摘要:

通过链式思考(CoT)使大型语言模型(LLMs)能够解决复杂任务,但因生成长理性而带来显著的推理成本。我们提出Thinking States方法,在输入处理过程中进行推理。具体而言,Thinking States每隔几个输入token生成一组思考token,将这些思考转换回嵌入空间,并将其添加到后续的输入token中。这种方法有两个关键优势。第一,它捕捉了CoT的递归性质,但思考token是在输入处理过程中生成的。第二,由于思考被表示为token,它们可以通过自然语言监督进行学习,并通过教师强制进行并行化学习。实验证明,Thinking States在多个推理任务中优于其他潜在推理方法,缩小了与CoT在数学问题上的差距,并在2跳问答任务中表现相当,但延迟更优。在状态跟踪任务中,我们展示了Thinking States比CoT具有更强的推理行为,能够成功外推到比训练期间看到更长的序列。

英文摘要:

Reasoning with a chain-of-thought (CoT) enables Large Language Models (LLMs) to solve complex tasks but incurs significant inference costs due to the generation of long rationales. We propose Thinking States, a method that performs reasoning {\em while} the input is processing. Specifically, Thinking States generates sequences of thinking tokens every few input tokens, transforms the thoughts back into embedding space, and adds them to the following input tokens. This has two key advantages. First, it captures the recurrent nature of CoT, but where the thought tokens are generated as input is processing. Second, since the thoughts are represented as tokens, they can be learned from natural language supervision, and using teacher-forcing, which is parallelizable. Empirically, Thinking States outperforms other latent reasoning methods on multiple reasoning tasks, narrowing the gap to CoT on math problems, and matching its performance on 2-Hop QA with improved latency. On state-tracking tasks, we show Thinking States leads to stronger reasoning behavior than CoT, successfully extrapolating to longer sequences than seen during training.

↑