arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

改进自适应循环Transformer的测试时扩展

Improving Test-Time Scaling with Adaptive Looped Transformers

Yichen You, Tianyu Fu, Aosong Feng, Xingtai Lv, Xuefei Ning, Ning Ding, Yu Wang

arXiv 2609.35748首次发表:更新:

发表机构

Tsinghua University; Yale University(清华大学; 耶鲁大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

提出TaH2,通过前瞻深度监督自适应分配循环迭代,提升循环Transformer的测试时扩展效率与准确率,在AIME基准上斜率提升53%。

AI 中文摘要

循环Transformer通过重用层进行潜在计算,展现了有前景的参数效率。先前的研究在匹配参数或每token FLOPs的条件下比较了循环和非循环模型。然而,据我们所知,随着输出变长,循环是否改善测试时扩展仍未得到充分探索。通过后训练循环Transformer,我们研究了准确率-计算量斜率,该斜率以每次测试时解码FLOPs加倍时的准确率增益来衡量。我们发现,现有的循环Transformer通常产生比非循环基线更陡的斜率,但在匹配计算量下表现不如基线。虽然固定深度循环对每个token都花费额外的迭代,但我们的分析表明,许多token并未从额外迭代中受益。因此,我们提出了TaH2,使模型能够将额外迭代集中在那些从循环中受益的token上。它通过前瞻深度监督联合后训练主干网络和迭代决策器,该监督使用在线标签指示进一步迭代是否改善预测。TaH2提高了测试时扩展的效率和可达到的准确率。在具有挑战性的AIME基准上,TaH2将准确率-计算量斜率比非循环基线提高了53%(2.74对比1.79),在匹配的测试时计算量下,超过基线峰值准确率约3.4个百分点。随着最大迭代深度的增加,现有循环模型大多趋于平稳,而TaH2相对于非循环基线的增益从深度2的+2.8个百分点持续增长到深度8的+3.9个百分点。我们的代码可在https URL获取。

英文摘要

Looped transformers have demonstrated promising parameter efficiency by reusing layers for latent computation. Prior studies compare looped and non-looped models at matched parameters or per-token FLOPs. However, to the best of our knowledge, whether looping improves test-time scaling as outputs grow longer remains underexplored. Through post-training looped transformers, we study the accuracy-compute slope, measured as the accuracy gain per doubling of test-time decoding FLOPs. We find that existing looped transformers often yield steeper slopes than their non-looped baseline, yet underperform it at matched compute. While fixed-depth looping spends extra iterations on every token, our analysis shows that many tokens do not benefit from extra iterations. We therefore propose TaH2, which enables the model to focus extra iterations on the tokens that benefit from looping. It jointly post-trains the backbone and an iteration decider through lookahead depth supervision, which uses online labels indicating whether further iteration improves the prediction. TaH2 improves both the efficiency and attainable accuracy of test-time scaling. On challenging AIME benchmarks, TaH2 improves the accuracy-compute slope by 53% (2.74 vs. 1.79) over the non-looped baseline, exceeding the baseline's peak accuracy by about 3.4 points at matched test-time compute. As the maximum iteration depth increases, existing looped models largely plateau, while TaH2's gain over the non-looped baseline continues to grow from +2.8 points at depth 2 to +3.9 points at depth 8. Our code is available at https://github.com/thu-nics/TaH.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑