发表机构
Princeton University; University of Pennsylvania(普林斯顿大学; 宾夕法尼亚大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
提出循环循环 Transformer(RLT),通过循环解码器使计算路径随序列长度增长,在算法任务上显著提升状态跟踪泛化能力,超越固定深度 Transformer。
AI 中文摘要
状态跟踪需要在每个输入处进行更新,但 Transformer 对每个 token 应用的深度是固定的,与序列长度无关。我们引入了循环循环 Transformer(RLT),它将层拆分在并行因果编码器和循环解码器之间。在每个 token 处,解码器将编码器输出与前一 token 的最终解码器状态合并,因此计算路径随序列长度增长,而每个 token 的成本固定。在六个算法任务上,我们比较了八层 Transformer 的五个拆分,每个拆分包含八层,跨越三个随机种子。在最多 40 位的训练下,两个 RLT 拆分将奇偶校验泛化到 256 位,每个种子准确率达到 100%,而 Transformer 仍停留在随机水平。在基于交换的 $S_5$ 置换跟踪中,在训练长度的八倍处,RLT 达到 97% 的最终状态准确率,而 Transformer 低于 1%,且准确率随解码器深度增加。在超出训练长度的模算术任务中,RLT 达到高达 93% 的准确率,而 Transformer 为 33%。消融实验表明,这些增益依赖于反馈:移除反馈会使奇偶校验和基于交换的 $S_5$ 在每个拆分上都降至随机水平。每四个 token 块更新一次反馈,允许块内已知 token 并行运行,并保持 64 位奇偶校验准确率为 99%,而置换跟踪依赖于逐 token 反馈:分块将长度 64 的基于交换的 $S_5$ 准确率从 100% 降至 20%。
英文摘要
State tracking requires an update at every input, but the depth a Transformer applies to each token is fixed regardless of sequence length. We introduce the Recurrent Looped Transformer (RLT), which splits its layers between a parallel causal encoder and a recurrent decoder. At each token, the decoder merges the encoder output with the previous token's final decoder state, so the computation path grows with sequence length at a fixed per-token cost. On six algorithmic tasks, we compare five splits of eight layers with an eight-layer Transformer over three seeds. Trained on at most 40 bits, two RLT splits generalize parity to 256 bits with 100% accuracy in every seed, while the Transformer stays at chance. On swap-based $S_5$ permutation tracking at eight times the training length, RLT reaches 97% final-state accuracy versus under 1% for the Transformer, and accuracy increases with decoder depth. On modular arithmetic beyond the training lengths, RLT reaches up to 93% versus 33% for the Transformer. Ablations show that these gains depend on the feedback: removing it drops parity and swap-based $S_5$ to chance at every split. Updating the feedback once per four-token chunk lets known tokens in a chunk run in parallel and keeps 64-bit parity at 99%, while permutation tracking depends on per-token feedback: chunking lowers length-64 swap-based $S_5$ from 100% to 20%.
CommentsProject Page: https://github.com/yifanzhang-pro/recurrent-looped-tranformer