arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.03624cs.CL

LoopMTP:由潜在多令牌预测引导的循环Transformer

LoopMTP: A looped transformer guided by latent multi-token prediction

Behzad Shomali, Markus Frey, David Berghaus, Joachim Koehler, Mehdi Ali

AI总结:

LoopMTP针对循环Transformer的潜在过度思考问题,通过潜在多令牌预测实现软对齐与轻量门控,使平均准确率提升最高8.1%,15次循环内训练稳定。

AI中文摘要:

循环Transformer已成为在固定参数规模下实现强推理能力的参数高效方案,其通过在T次迭代中复用一层网络,获得了更大模型的有效深度与推理能力。然而现有方法存在潜在过度思考与计算无差异化问题,主要原因是中间表示在各循环间缺乏指导。多令牌预测(MTP)恰好提供了循环所缺失的密集前瞻性监督。我们提出的LoopMTP通过潜在空间的结构关联将两者结合:循环T次的模型可预测T个未来令牌,具体通过将第t次循环的隐状态与t步后令牌的嵌入进行软对齐,同时用轻量门控在迭代间保留有用信息。LoopMTP相较于非循环基准,平均准确率提升最高达8.1%(相对),且在多达15次循环时训练仍保持稳定。

英文摘要:

Looped transformers have emerged as a parameter-efficient alternative to scaling depth for strong reasoning. By reusing one stack of layers across $T$ iterations, they attain the effective depth and reasoning capabilities of larger models at a fixed parameter count. Yet existing approaches suffer from latent overthinking and undifferentiated computation, largely because intermediate representations receive no guidance across loops. Multi-token prediction (MTP) supplies exactly the dense, forward-looking supervision the loop is missing. We propose \textsc{LoopMTP}, which links the two through a structural correspondence in latent space: a model that loops $T$ times can anticipate $T$ future tokens. \textsc{LoopMTP} realizes this by softly aligning the hidden state of loop $t$ with the embedding of the token $t$ steps ahead, while a lightweight gate preserves useful information across iterations. \textsc{LoopMTP} improves average accuracy by up to 8.1\% (relative) over the non-looped baseline, with training remaining stable for up to 15 loops.

↑