arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.21605cs.LG

在循环Transformer中用时间换取深度

Trading Depth for Time in Recurrent Transformers

发表机构微软 · 威斯康星大学麦迪逊分校
查看机构详情
  • Microsoft(微软)
  • University of Wisconsin–Madison(威斯康星大学麦迪逊分校)

机构由 AI 辅助整理,请以论文原文为准。

Zeyi Huang, Xuehai He, Yong Jae Lee, Yelong Shen

首次发表
浏览论文内容

中文总结 AI 辅助

本研究通过潜在循环Transformer比较时间递归与物理深度增加,发现插入思维词元可在减少约48%参数的情况下恢复双深度模型67%-81%的性能改进,表明时间思考是参数高效的深度替代方案。

中文摘要 AI 辅助

循环Transformer通过时间递归增加计算深度,将每个词元的高层隐藏状态馈送到下一个词元的计算中。这引发了一个自然的问题:额外的计算是更好地用于更多的时间步,还是更大的物理深度?我们使用潜在循环Transformer(LRTs)来研究这个问题,LRTs在解码期间每个词汇词元保留一次骨干网络前向传播,并为比较这两种增加计算的方式提供了受控设置。具体来说,我们在连续的词汇词元之间插入一个潜在思维词元。每个思维词元与词汇词元一样通过相同的$L$层,共享骨干网络参数,并在预测下一个词元之前提供额外的隐藏状态细化阶段。我们将这个$L$层LRT与没有思维词元的$2L$层LRT进行比较。两者在解码期间每个词汇词元执行$2L$个Transformer块,但思维词元模型使用更少的参数。在16层和20层的混合专家NanoChat骨干网络上,一个思维词元使较浅模型在每字节比特数上分别达到其双深度对应模型的0.006和0.004以内,恢复了67%和81%的改进,同时总参数减少约48%。这些结果表明,时间思考为增加循环Transformer的物理深度提供了一种参数高效的替代方案。

英文摘要

Recurrent Transformers increase computational depth through temporal recurrence, feeding each token's high-level hidden state into the computation of the next. This raises a natural question: is additional computation better spent on more temporal steps or greater physical depth? We investigate this question using Latent Recurrent Transformers (LRTs), which retain one backbone forward pass per vocabulary token during decoding and provide a controlled setting for comparing these two ways of adding computation. Specifically, we insert a latent thought token between consecutive vocabulary tokens. Each thought token passes through the same $L$ layers as a vocabulary token, sharing the backbone parameters and providing an additional stage of hidden-state refinement before predicting the next token. We compare this $L$-layer LRT against a $2L$-layer LRT without thought tokens. Both execute $2L$ Transformer blocks per vocabulary token during decoding, but the thought-token model uses fewer parameters. On 16- and 20-layer mixture-of-experts NanoChat backbones, one thought token brings the shallower model within 0.006 and 0.004 bits per byte of its double-depth counterpart, recovering 67% and 81% of the improvement with approximately 48% fewer total parameters. These results suggest that temporal thinking offers a parameter-efficient alternative to increasing physical depth in recurrent Transformers.

↑