发表机构
IBM Research; Cornell University(IBM研究院; 康奈尔大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
循环Transformer中共享内存可提升质量并大幅减少内存使用,新模型LPT及其混合变体在FineWeb-Edu上以更少内存取得更低困惑度。
AI 中文摘要
循环Transformer对每个token多次应用相同的层,通过增加计算量来提升质量,而不增加参数数量。然而,每次递归都会写入自己的键值缓存,因此内存仍然随计算量增长。推理时技术可以缩小此缓存,但会牺牲质量。我们预训练循环语言模型以共享内存:仅第一次递归写入缓存,后续递归读取该缓存,同时保留自身的一个短窗口。令人惊讶的是,我们发现共享内存不会损害质量,反而会提升质量。在150M-1B参数规模下,我们的循环预测Transformer(LPT)及其混合变体为循环模型设立了新的质量-内存前沿:在五次递归下,混合变体在FineWeb-Edu上将验证困惑度相对于同尺寸标准Transformer降低了1.12-1.82,同时使用了76-79%更少的上下文内存。通过广泛的分析,我们探究了内存共享为何有帮助。共享内存和局部内存发展出不同的表示,后续递归主要关注共享内存,而共享内存也充当了通向第一次递归的梯度高速公路。
英文摘要
Looped Transformers apply the same layers several times per token, adding compute to improve quality without more parameters. Each recursion, however, writes its own key-value cache, so memory still grows with compute. Inference-time techniques can shrink this cache at a cost in quality. We pretrain looped language models to share memory: only the first recursion writes a cache, and later recursions read it while keeping a short window of their own. Surprisingly, we find that sharing memory does not cost quality and instead improves it. At 150M-1B parameters, our Looped Prediction Transformer (LPT) and its hybrid variant set a new quality-memory frontier for looped models: with five recursions, the hybrid lowers validation perplexity on FineWeb-Edu by 1.12-1.82 relative to a same-size standard Transformer while using 76-79% less context memory. Through an extensive analysis, we investigate why memory sharing helps. Shared and local memory develop different representations, and later recursions attend mostly to the shared memory, which also acts as a gradient highway to the first recursion.