重新思考测试时训练的表达性与效率
Rethinking Expressivity and Efficiency in Test-Time Training
- Fraunhofer IOSB(弗劳恩霍夫IOSB研究所)
- National University of Singapore(新加坡国立大学)
- Lamarr Institute for Machine Learning and Artificial Intelligence(拉马尔机器学习与人工智能研究所)
- University of Bonn(波恩大学)
- Karlsruhe Institute of Technology (KIT)(卡尔斯鲁厄理工学院)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
该研究针对测试时训练难以平衡表达性与效率的问题,提出E²-TTT方法,通过推导闭式状态转移实现并行分块训练,在语言建模、上下文检索及长上下文外推任务中表现优异,兼顾性能与训练效率。
AI中文摘要:
测试时训练(Test-Time Training, TTT)通过在推理过程中进行连续权重更新实现长上下文处理,但现有方法难以平衡每token更新动态的表达性与分块近似的硬件效率。我们提出E²-TTT(Expressive and Efficient TTT,兼具表达性与效率的测试时训练)以弥合这一差距。在采用分块起始权重计算梯度的标准近似下,我们推导了闭式状态转移,其可精确复现每token循环的分块末端快速权重与动量状态。这使得分块级训练可完全并行化,同时保留了此前分块方法所丢弃的更新规则时间结构。我们通过从头训练参数规模达13亿(1.3B)的模型验证了E²-TTT的有效性。在语言建模任务中,其性能与现有TTT及混合注意力基线相当,而在上下文检索任务中优于这些基线。其优势在长度外推场景中最为显著:在标准“干草堆中的针”(Needle in a Haystack)密码测试中,当上下文长度为训练长度的8倍时,它仍保持90%以上的准确率。同时,E²-TTT可达到高效分块方法的训练吞吐量,证明其能有效协调表达性与效率。代码可从指定URL获取。
英文摘要:
Test-Time Training (TTT) enables long-context processing via continuous weight updates during inference, but current methods struggle to balance the expressivity of per-token update dynamics with the hardware efficiency of chunk-wise approximations. We propose E$^2$-TTT (Expressive and Efficient TTT) to bridge this gap. Under the standard approximation of taking gradients at the chunk-start weights, we derive a closed-form state transition that exactly reproduces the chunk-end fast-weight and momentum states of the per-token recurrence. This enables fully parallelized chunk-level training while preserving the temporal structure of the update rule that prior chunk-wise methods discard. We validate E$^2$-TTT by training models up to 1.3B parameters from scratch. It performs on par with previous TTT and hybrid attention baselines in language modeling while outperforming them on in-context retrieval. Its advantage is most pronounced in length extrapolation: on the standard ``Needle in a Haystack'' passkey test, it retains over 90% accuracy at $8\times$ the training context length. Meanwhile, E$^2$-TTT can match the training throughput of efficient chunk-wise methods, demonstrating that it effectively reconciles expressivity with efficiency. The code is available at https://github.com/zeyun-zhong/E2-TTT.