arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

什么限制了递归推理模型:优化、架构与测试时扩展

What Limits Recursive Reasoning Models: Optimization, Architecture and Test-Time Scaling

Yuliana Shakhvalieva, Dmitrii Kharchev, Viacheslav Bezrukov, Inessa Fedorova, Dmitry Bocharov, Ivan Oseledets, Valerii Ternovskii

arXiv 2609.39967首次发表:更新:

发表机构

RND NLP, DAIMLD(RND NLP,DAIMLD)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究通过统一实验揭示递归推理模型性能的关键在于优化策略而非架构,提出稳定训练方案,构建13.6M参数模型,在算法任务上显著超越现有基线。

AI 中文摘要

递归推理模型多次应用一个共享的小型Transformer块来精炼潜在状态。这使它们以少量参数获得较大的有效深度,并在算法任务上表现强劲。这类紧凑求解器是LLM在狭窄算法子问题上可调用的工具的自然候选。然而,现有模型如HRM、TRM和URM在架构、梯度传播和训练过程上同时存在差异,这使得难以判断其性能的驱动因素,且其优化仍未被充分理解且常不稳定。在本工作中,我们同时解决这两个空白。首先,我们在涵盖六个算法领域的统一实验流程下研究这些问题。在代表性领域上执行个体受控消融,而所得方案在完整套件上评估。研究揭示了一个令人惊讶的简单且可泛化的递归推理方案:中间梯度视界、大批量物理批次和受控的递归状态更新。显式的分层架构并非必需。其次,我们将这些发现结合成一个稳定的13.6M参数模型,在评估的递归基线中实现了最强的整体性能,尤其在分布外泛化上取得显著提升。它将Arithmetic OOD准确率从最强基线的36.2%提升至71.2%,同时在Sudoku上达到98.41%,在ARC-AGI-1上达到59.5%的pass@2。我们的结果表明,在本文研究的递归架构中,性能强烈依赖于递归的优化和稳定化方式。更广泛地,它展示了如何通过逐一优化组件来改进AI系统。

英文摘要

Recursive reasoning models apply a small shared Transformer block many times to refine a latent state. This gives them large effective depth with few parameters and makes them strong on algorithmic tasks. Such compact solvers are natural candidates for tools that an LLM can call on narrow algorithmic subproblems. However, existing models such as HRM, TRM and URM differ in architecture, gradient propagation and training procedure simultaneously. This makes it hard to tell what drives their performance, and their optimization is still poorly understood and often unstable. In this work we address both of these gaps. First, we study these questions under a unified experimental pipeline spanning six algorithmic domains. Individual controlled ablations are performed on representative domains, while the resulting recipe is evaluated across the full suite. The study reveals a surprisingly simple recipe for stable and generalizable recursive reasoning: an intermediate gradient horizon, large physical batches and controlled updates of the recurrent state. An explicit hierarchical architecture is not needed. Second, we combine these findings into a stable 13.6M-parameter model that achieves the strongest overall performance among the evaluated recursive baselines, with particularly large gains on out-of-distribution generalization. It raises Arithmetic OOD accuracy to 71.2%, from 36.2% for the strongest baseline, while reaching 98.41% on Sudoku and 59.5% pass@2 on ARC-AGI-1. Our results show that, within the recursive architectures studied here, performance depends strongly on how recurrence is optimized and stabilized. More broadly, it shows how AI systems can be improved by optimizing their components one at a time.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑