arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

T^2MLR:具有时间中层循环的Transformer

T^2MLR: Transformer with Temporal Middle-Layer Recurrence

Ziyang Cai, Xingyu Zhu, Yihe Dong, Yinghui He, Sanjeev Arora

arXiv 2607.15178首次发表:更新:

发表机构

Princeton University(普林斯顿大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究针对Transformer推理受自回归解码限制问题,提出T2MLR架构,将前token缓存中间层表示融合到当前token较早层,在自然语言预训练和多跳推理微调中表现优,且局部中层块应用循环效果好,还无需从头预训练,降低实际应用门槛。

AI 中文摘要

Transformer推理受自回归解码限制,通过token空间反复压缩丰富的隐藏计算,使中间推理状态难以持久。我们引入了具有时间中层循环(T2MLR)的Transformer,将前一个token的缓存中间层表示直接融合到当前token位置的较早层,使抽象中间计算在解码步骤中持久,推理开销小。在自然语言预训练和多跳推理微调中,T2MLR始终优于基线。仅对局部中层块应用循环(低至网络的20%)往往优于全层循环。T2MLR无需从头预训练,将循环路径改造到现有预训练的17亿参数Transformer并微调可显著提升数学推理。结果表明有效的Transformer潜在推理无需像以前那样遍历所有层,通过有针对性的中层循环即可更有效实现。

英文摘要

Transformer reasoning is limited by autoregressive decoding, which repeat edly compresses rich hidden computation through token space and makes it difficult for intermediate reasoning states to persist across time. We in troduce Transformers with Temporal Middle-Layer Recurrence (T2MLR), a transformers-based latent reasoning architecture that fuses a cached middle layer representation from the previous token directly into an earlier layer of the current token position, enabling abstract intermediate computation to persist across decoding steps with little inference overhead. Across natural-language pretraining and multi-hop reasoning finetuning, T2MLR consistently outperforms data- and parameter-matched Transformer base lines. Moreover, applying recurrence to only a localized middle-layer block (as little as 20% of the network) often outperforms full-layer recurrence. Im portantly, T2MLR does not require pretraining from scratch: retrofitting the recurrent pathway into an existing pretrained 1.7B Transformer and briefly finetuning substantially improves math reasoning, lowering the barrier to practical adoption. These results suggest that effective latent reasoning in Transformers does not require looping over all layers as in previous works, but can instead emerge more strongly from targeted middle-layer recurrence.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑