arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.15160cs.AI

T-LoopFormer:用于潜在推理的动态路由令牌级弹性深度循环Transformer

T-LoopFormer: Token-Level Elastic-Depth Looped Transformers for Latent Reasoning with Dynamic Routing

Mingqian Yu, Wenpeng Zhang, Shaobo Cui, Peilin Zhao

首次发表
浏览论文内容

中文总结 AI 辅助

T-LoopFormer通过动态令牌选择路由和递归级KV缓存,实现令牌级弹性深度循环Transformer,在潜在推理中优化计算分配,达到最先进性能并降低推理延迟。

中文摘要 AI 辅助

循环Transformer最近通过在多次迭代中重用一组共享参数,在推理和语言任务中均展现出强大性能,实现了参数效率而不牺牲表示能力。此外,循环Transformer直接在潜在空间中进行推理(潜在推理),以减少推理期间消耗的令牌数量,从而提高样本效率。然而,这些模型通常对所有令牌统一应用固定的递归深度,导致计算分配次优,并留下了显著的效率提升空间。在这项工作中,我们提出了循环Transformer的动态令牌选择路由,使每个令牌能够根据其隐藏状态自适应地确定其自身的循环迭代次数。我们使用动态路由器来决定令牌是应继续递归还是提前退出,允许简单令牌绕过不必要的计算,而困难令牌则获得更深入的处理。为确保这种自适应机制不损害解码效率,我们进一步引入了递归级KV缓存,该缓存为每个递归循环维护独立的键值缓存。这种设计确保不同深度的令牌仅关注其对应的缓存状态,有效消除了已退出令牌的冗余计算,并实现了快速自回归解码。大量实验表明,在相同参数下,T-LoopFormer在PPL和10个零样本推理任务上达到了最先进的性能,甚至在24倍FLOPs下超越了基础模型,且我们的模型能达到最低的推理延迟,这验证了令牌选择路由器和递归级KV缓存的有效性。代码:此https URL。

英文摘要

Looped Transformers have recently demonstrated strong performance in both reasoning and language tasks by reusing a shared set of parameters across multiple iterations, achieving parameter efficiency without sacrificing representational power. Besides, looped Transformers perform inference directly in the latent space (latent reasoning) to reduce the number of tokens consumed during inference, thereby achieving improved sample efficiency. However, these models typically apply a fixed recursion depth uniformly to every token, leading to suboptimal compute allocation and leaving significant efficiency gains on the table. In this work, we propose dynamic token-choice routing for looped transformers, enabling each token to adaptively determine its own number of loop iterations based on its hidden state, which can improve the token generation accuracy. Moreover, we further introduce recursion-wise KV cache, which maintains an independent key-value cache for each recursion loop, this design ensures that tokens at different depths only attend to their corresponding cached states, effectively enabling faster autoregressive decoding. Extensive experiments show that T-LoopFormer achieves robust performance on language modeling and zero-shot reasoning tasks and our model can reach the lowest decoding latency, which validate the effectiveness of token-choice router and recursion-wise KV cache. Code is available at https://github.com/YuMingQian1234/T-LoopFormer

↑