arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.20594cs.LGcs.AIstat.ML

循环何时成为一种算法?权重绑定循环变换器中的收敛选择

When Does Recurrence Become an Algorithm? Convergence Selection in Weight-Tied Looped Transformers

Tong Zhang, Junhao Hu, Yun Peng, Tao Xie

首次发表
浏览论文内容

中文总结 AI 辅助

研究权重绑定循环变换器何时实现算法,通过群字问题得出预算定律、架构先验决定算法等四个发现,介绍收敛时间缩放工具并验证其因果关系,结果可复现。

中文摘要 AI 辅助

我们探讨了权重绑定循环变换器(一个模块应用T次)何时能实现实际算法。通过对群字问题的受控总体研究,得出四个发现。一是预算定律,自由训练确定线性计算前沿,SGD选择符合要求的前沿,据此有原则性的停止规则。二是架构先验而非表达能力决定算法选择。三是电路复杂度与实际情况不符,特定算子课程可解决问题。四是机制具有可移植性而非强制性。还介绍了新工具并验证其因果关系,结果可在公共基准上复现。

英文摘要

When does a weight-tied looped transformer -- one block applied T times -- implement an actual algorithm? We answer with four findings from controlled populations on group word problems. (1) The budget law: free training installs a linear computation frontier, a mechanism that solves v positions per loop, whose speed is priced by the training contract: v ~ n_train/T_train (exponent 0.98 +/- 0.04, R^2=0.99), exactly unity under T=n training. SGD selects a frontier matching the minimum the contract demands; granting more test-time loops than ever trained rescues late positions at fixed input length, yielding a principled halting rule T* = ceil(n / v-hat). (2) Architecture prior, not expressivity, picks the algorithm: standard-depth transformers learn parallel scans on this family; weight tying flips the selection to the serial frontier, even when positional addressing for a log-depth scan is supplied. At matched depth and parameters, untied models extrapolate worst and fail to learn A5 at all. (3) The walls are not where circuit complexity says: NC1-completeness costs nothing (A5 generalizes fully), while group order does (S5's 120x120 operator deadlocks joint learning) -- and an operator-first curriculum dissolves the wall in every seed. (4) Mechanisms are portable, not mandatable: warm-starting across budget contracts transfers the algorithm in every seed, re-pricing its speed, while imposing seriality through the input schedule fails where free training succeeds. These results are invisible to standard instruments, which provably saturate at the fixed points trained loops converge to. We introduce a head instrument, the convergence-time scaling tau(n,i), validate it causally via damage cones whose slope reproduces v, and show in-distribution head measurements predict out-of-distribution fate where tail metrics do not. Results replicate on the public easy-to-hard benchmark.

发表机构

  • Fudan University(复旦大学)
  • Peking University(北京大学)

机构由 AI 辅助整理,请以论文原文为准。

↑