arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.38149cs.CL

预训练带教师监督的潜在信息反馈Transformer

Pretraining Latent Information Feedback Transformers with Teacher Supervision

Dor Tirosh, Ido Amos, Mor Geva

首次发表
浏览论文内容

中文总结 AI 辅助

本研究提出LIFT架构,通过教师监督将循环状态学习转为预测问题,使Transformer在预训练时实现深层到浅层反馈,在多项任务上优于标准模型。

中文摘要 AI 辅助

Transformer语言模型(LMs)是前馈式的:深层表示从不反馈到较浅层,跨生成步骤信息向下流动的唯一途径是解码后的token。这种狭窄的通道迫使模型重新计算中间结果并丢弃备选续写。在这项工作中,我们在预训练期间移除这一瓶颈,引入了LIFT(潜在信息反馈Transformer)架构和训练方法,使语言模型能够跨生成步骤传播状态。我们通过将循环状态学习转化为教师强制预测问题来实现这一点:每个输入token与一个信息密集的状态配对,该状态源自现成的预训练语言模型的下一个token分布。模型通过增加少量额外参数进行扩展,然后被训练以同时预测下一个token和下一个状态。由于输入状态是预先计算的,预训练在位置间保持完全并行。在推理时,模型自身预测的状态被反馈,计算开销很小且随模型规模增大而减小。使用从135M到1B参数的预训练模型进行的实验表明,在token匹配预算下,LIFT在语言建模、下游推理任务和程序任务上始终优于标准Transformer和基线,同时与计算匹配的Transformer相当或领先。此外,在状态跟踪任务上的受控研究表明,即使使用在任务上失败的Transformer的状态进行训练,一个微小的LIFT也优于在8倍数据上训练的相同规模Transformer。总体而言,我们表明语言模型可以通过可扩展的教师监督在预训练期间学习利用从深层到浅层的反馈。

英文摘要

Transformer language models (LMs) are feed-forward: deep-layer representations are never fed back to shallower layers, and the only pathway for information to flow downward across generation steps is the decoded token. This narrow channel forces models to recompute intermediate results and to discard alternative continuations. In this work, we remove this bottleneck during pretraining, introducing the LIFT (Latent Information Feedback Transformer) architecture and training method which enable LMs to propagate state across generation. We achieve this by turning recurrent-state learning into a teacher-forced prediction problem: each input token is paired with an information-dense state, derived from the next-token distribution of an off-the-shelf pretrained LM. The model, extended with a small number of additional parameters, is then trained to predict both the next token and the next state. As the input states are precomputed, pretraining remains fully parallel across positions. At inference, the model's own predicted states are fed back, with a minor computational overhead that decreases with model size. Experiments with pretrained models ranging from 135M to 1B parameters show that LIFT consistently outperforms standard Transformers and baselines on language modeling, downstream reasoning tasks, and procedural tasks under token-matched budget, while being on par with or ahead of compute-matched Transformers. Moreover, a controlled study on a state-tracking task shows that a tiny LIFT outperforms same-size Transformers trained on 8x more data, even when trained with the states of a Transformer that fails the task. Overall, we show that LMs can learn to exploit deep-to-shallow feedback during pretraining via scalable teacher supervision.

发表机构

  • Blavatnik School of Computer Science and AI, Tel Aviv University(特拉维夫大学布拉瓦特尼克计算机科学与人工智能学院)
  • The Hebrew University of Jerusalem(耶路撒冷希伯来大学)

机构由 AI 辅助整理,请以论文原文为准。

↑