arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

当递归模型完成计算时

When Recursive Models Finish Computing

Hare Krishna, Shubham Singh, Stephen Ebert, Hao-Yu Sun

arXiv 2609.26487首次发表:更新:

发表机构

University of Texas at Austin; Tufts University; Zyphra Technologies; Austin Community College; SpaceXAI(德克萨斯大学奥斯汀分校; 塔夫茨大学; Zyphra 科技公司; 奥斯汀社区学院; SpaceXAI 公司)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究通过扩展递归步数,发现注意力与MLP递归模型在数独任务上完成计算具有轨迹条件各向异性稳定性,区分了预算失败与计算完成。

AI 中文摘要

递归模型可以在其名义推理预算之外继续更新其潜在状态,因此在预算内产生的错误输出并不能表明计算是未完成还是已进入持续失败的阶段。我们研究了基于注意力和MLP的微型递归模型(TRMs)在1,000个困难数独谜题上的完成动态。将递归从名义上的16步扩展到512步,注意力模型的累积精确求解准确率从59.2%提高到87.5%,MLP模型从74.4%提高到91.9%,解决了名义预算中未解决谜题的三分之二以上。在两种架构中,潜在状态运动在首次精确解之后急剧下降。完成状态通常沿轨迹方向是局部收缩的,即使相同的局部雅可比矩阵保留了强扩张方向。我们将这种现象表征为轨迹条件各向异性稳定性。扰动实验证实了两种模型的这种方向稳定性。最大扩张方向的多步命运不同:在注意力模型中,它在16步内被吸收,但在MLP模型中持续更长时间。各向异性稳定性模式也适用于第二个注意力检查点。总之,这些结果区分了名义预算失败与完成计算,并识别了两种递归架构中完成的共同动态特征。

英文摘要

Recursive models can continue updating their latent states beyond their nominal inference budget, so an incorrect output at that budget does not show whether computation is unfinished or has entered a persistently unsuccessful regime. We study the dynamics of completion in attention- and MLP-based Tiny Recursive Models (TRMs) on 1,000 hard Sudoku puzzles. Extending recurrence from the nominal 16 steps to 512 steps increases cumulative exact-solve accuracy from 59.2% to 87.5% for the attention model and from 74.4% to 91.9% for the MLP model, solving more than two-thirds of the puzzles unsolved in the nominal budget. Across both architectures, latent-state motion drops sharply after the first exact solution. Completed states are typically locally contractive along the trajectory direction, even though the same local Jacobian retains strongly expanding directions. We characterize this phenomenon as trajectory-conditioned anisotropic stability. Perturbation experiments confirm this directional stability across both models. The multi-step fate of the maximally expanding direction differs: it is absorbed within 16 steps in the attention model but persists longer in the MLP model. The anisotropic-stability pattern also holds for a second attention checkpoint. Together, these results distinguish nominal-budget failure from completed computation and identify a common dynamical signature of completion across two recurrent architectures.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑