arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

循环变换器中的自适应深度:诊断学习到的停止门和轨迹读出

Adaptive Depth in Looped Transformers: Diagnosing Learned Halting Gates and Trajectory Readouts

Andrei Cristian Popescu, Haitz Sáez de Ocáriz Borde, Pietro Liò

arXiv 2607.20519首次发表:更新:

发表机构

University of Cambridge(剑桥大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究循环变换器中自适应深度,通过轨迹——读出视角,涵盖合成任务与大规模检查点。发现固定先验深度监督可产生有益轨迹,简单读出表现良好,拟合冻结轨迹定位失败原因,Ouro评估有类似情况,重新定义自适应深度为联合问题。

AI 中文摘要

循环变换器通过重复应用共享循环块增加测试时计算量。循环变换器中学习到的停止目标通常使用单一退出分布作为推理时停止规则和训练时每层损失加权。这使退出选择与轨迹形成纠缠,导致自适应计算性能不佳。我们通过轨迹——读出视角研究循环变换器中的自适应深度,涵盖合成任务及大规模检查点。发现固定先验深度监督能产生困难感知轨迹,简单事后置信度读出常匹配或优于学习到的线性和多层感知器门。拟合冻结轨迹上的门定位失败原因,在Ouro评估中也有相同模式,预训练思考门有竞争力但非均匀帕累托最优,测量延迟表明平均退出深度减少转化为实际推理时间节省。我们的系统诊断评估将循环变换器中的自适应深度重新定义为轨迹形成和退出读出的联合问题,而非仅门学习问题,突出了先前学习停止工作常隐含的区别。

英文摘要

Looped Transformers increase test-time computation by repeatedly applying a shared recurrent block. Learned halting objectives in looped Transformers typically use a single exit distribution both as the inference-time stopping rule and as the training-time weighting of per-depth losses. This entangles exit selection with trajectory formation: the gate not only chooses which recurrent state to use, but also determines how strongly each intermediate state is supervised. Consequently, poor adaptive-compute performance can arise from the readout, the induced trajectory, or their interaction. We study adaptive depth in looped Transformers through this trajectory--readout lens, across controlled synthetic tasks (modular arithmetic and binary parity) and large-scale Ouro-1.4B and 2.6B checkpoints. We find that fixed-prior depth supervision, which shapes the trajectory without an input-dependent halting policy, produces difficulty-aware trajectories whose intermediate states expose useful stopping signals, and that simple post-hoc confidence readouts often match or outperform learned linear and MLP gates. Fitting gates on frozen trajectories localizes the failure: it appears to stem mainly from the trajectory induced by joint gate training rather than from limited gate expressivity. The same pattern is present in Ouro evaluations, where pretrained ponder gates are competitive but not uniformly Pareto-optimal, and measured latency confirms that the resulting reductions in average exit depth translate into practical inference-time savings. Our systematic diagnostic evaluation reframes adaptive depth in looped Transformers as a joint problem of trajectory formation and exit readout, rather than gate learning alone, highlighting a distinction that prior learned-halting work has often left implicit.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑