arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

潜在循环思维:使用冻结大语言模型进行推理时对所提潜在表示的循环优化

Latent Recurrent Thoughts: Recurrent Refinement of Proposed Latents for Reasoning with Frozen LLMs

Zhaoliang Chen, Jie Fu

arXiv 2609.01117首次发表:更新:

发表机构

Emory University; IQuest Research(埃默里大学; 艾奎斯特研究院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究提出 LRT 方法,通过冻结 LLM 结合循环优化的潜在表示进行推理,在多个推理任务上优于同类冻结解码器方法及思维链提示,且推理计算量更低。

AI 中文摘要

思维链推理在离散的 token 空间中展开:每一步都以文本形式确定下来,错误会发生传播,而引出良好的推理轨迹则预设了存在可模仿的轨迹。而在模型的连续表示空间中进行推理——中间状态是向量而非单词——可以规避这些限制,但仍存在如何计算这些潜在状态的问题。我们从两个维度着手解决该问题:其一,我们保留大语言模型(LLM)为冻结状态,将其用于自身擅长的任务——序列建模与解码,同时由一个小型辅助网络提供连续的潜在思维作为输入;其二,我们通过循环方式生成这些潜在表示:一个小型循环推理器会在多步过程中对其进行优化,将计算深度与模型规模解耦,使潜在表示成为迭代处理的产物而非单次前向传播的结果。我们将此实例化为潜在循环思维(Latent Recurrent Thoughts, LRT):一个任务专用的提议器提供基础潜在表示,循环推理器通过有界残差修正对其进行优化,冻结的 LLM 则解码出答案。在带有答案监督但无推理轨迹的符号推理任务(Countdown-4、数独)以及自然语言推理任务(HumanEval、MBPP、StrategyQA)上,在相同的解码器、提示、数据和训练预算下,LRT 的性能显著优于先前的冻结解码器连续空间推理方法,且在相同主干模型上,其推理计算量仅为非思维模式思维链提示的一小部分,性能却优于后者。

英文摘要

Chain-of-thought reasoning unfolds in discrete token space: each step is committed as text, errors propagate, and eliciting good traces presupposes traces to imitate. Reasoning instead in a model's continuous representation space - where intermediate states are vectors rather than words - sidesteps these constraints, but leaves open how those latent states should be computed. We approach this along two axes. First, we keep a large language model (LLM) frozen and use it for what it is already good at - modeling and decoding sequences - while a small auxiliary network supplies continuous latent thoughts as input. Second, we produce those latents by recurrence: a tiny recurrent reasoner refines them over many steps, decoupling the depth of computation from the size of the model, so that the latents are a product of iterative processing rather than a single forward pass. We instantiate this as Latent Recurrent Thoughts (LRT): a task-dedicated proposer supplies base latents, a recurrent reasoner refines them through bounded residual corrections, and the frozen LLM decodes the answer. On symbolic reasoning with answer supervision but no reasoning traces (Countdown-4, Sudoku) and on natural-language reasoning (HumanEval, MBPP, StrategyQA), LRT substantially outperforms prior frozen-decoder continuous-space reasoning methods under an identical decoder, prompt, data, and training budget, and outperforms non-thinking-mode chain-of-thought prompting on the same backbone at a small fraction of its inference compute.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑