arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

RecurTrace:基于循环时间记忆的自适应潜在推理

RecurTrace: Adaptive Latent Reasoning with Loop-Time Memory

Yuxiang Wang, Kunyu Feng, Yingda Shen, Haoning Xu, Junyu Wang, Zhizheng Wu

arXiv 2609.03379首次发表:更新:

AI 中文总结

RecurTrace通过循环记忆注意力和自适应停止机制解决现有潜在循环推理的局限,在MathQA任务上优于多个基线,且在多参数规模下提升了生成准确率。

AI 中文摘要

重复一小部分中间层可在不增加参数或生成额外token的情况下提升语言模型的有效推理深度,近期研究表明这种潜在循环能改善推理能力。然而,两项设计选择限制了这些提升:每次迭代仅能看到前一次的输出,无法直接访问更早的计算结果;固定循环次数会在简单输入上浪费深度,而给困难输入的计算量不足。我们提出RecurTrace,利用循环自身轨迹解决这两个限制:具体而言,循环记忆注意力让每个循环层沿循环时间轴关注自身之前迭代的状态,使模型可重新访问更早的计算结果,而非仅依赖最新状态;随后,一个停止头读取循环状态并预测是否继续,由一个神谕(oracle)监督,该神谕会识别何时增加深度仍能降低损失。在相同循环骨干网络的受控MathQA对比中,RecurTrace以平均2.0次循环达到56.9%的准确率,在计算量匹配的情况下,比最优固定循环深度的结果高出2.2个百分点;相比之下,ACT和PonderNet均退化为1次循环,CALM在5.6次循环时仅达到54.1%,而更强的基线LoopUS-Conf和TaH-Mismatch分别在3.2次循环和2.1次循环时达到55.3%和55.7%。最后,RecurTrace在0.6B、1.7B、4B和8B参数规模下,相较于相同预算微调的基线,提升了生成准确率,且增益随模型规模增大而增加,从0.6个百分点升至3.4个百分点。

英文摘要

Repeating a small block of middle layers increases a language model's effective inference depth without adding parameters or generating extra tokens, and recent work shows that this latent recurrence improves reasoning. However, two design choices limit these gains. Each iteration sees only the previous output and cannot directly access earlier computations. Moreover, a fixed loop count wastes depth on easy inputs while leaving hard ones with too little computation. We introduce RecurTrace, which addresses both limitations using the loop's own trajectory. Specifically, Loop Memory Attention lets each looped layer attend to its own states from previous iterations along the loop-time axis, so the model can revisit earlier computations instead of relying on the latest state alone. A halting head then reads the loop state and predicts whether to continue, with supervision from an oracle that identifies when additional depth still reduces loss. In a controlled MathQA comparison on the same looped backbone, RecurTrace achieves 56.9% accuracy with an average of 2.0 loops, exceeding the best fixed loop depth by 2.2 points at matched compute. By comparison, ACT and PonderNet collapse to one loop, and CALM reaches only 54.1% with 5.6 loops, while the stronger LoopUS-Conf and TaH-Mismatch baselines reach 55.3% at 3.2 loops and 55.7% at 2.1 loops. Finally, RecurTrace improves generation accuracy over same-budget fine-tuned baselines at 0.6B, 1.7B, 4B, and 8B, with the gain growing with model size from 0.6 to 3.4 points.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑