arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.07110cs.LGcs.CL

模块化测试时训练(TTT):将测试时训练重新思考为可组合模块

Modular TTT: Rethinking Test-Time Training as Composable Modules

Bohao Tang, Zhen Qin, Yuqi Pan, Zheng Li, Pengfei Liu, Ya Zhang

首次发表
浏览论文内容

中文总结 AI 辅助

该研究提出Modular TTT框架,将TTT各组件设为显式设计维度,通过消融实验明确关键组件对性能的影响,训练的大参数模型性能可与Gated DeltaNet媲美。

中文摘要 AI 辅助

测试时训练(TTT)将序列建模视为在线学习问题,其中快速权重通过内部学习规则更新。尽管TTT变体数量不断增加,但现有方法通常单独对每个变体进行硬编码,这使得设计新的TTT方法以及隔离每个组件的作用变得困难。为解决该问题,我们提出Modular TTT,这是一个将内部学习者表示为有向无环图的框架,将快速权重网络、损失函数、学习率、权重衰减和归一化作为显式设计维度。Modular TTT自动将原语级别的训练视图前向、训练视图后向和因果查询视图规则组合为完整的图级TTT计算,包括快速权重状态转换。使用Modular TTT,我们系统地对TTT的组件进行消融实验,发现较小的学习率初始化、权重衰减和单层非线性可提升性能,而MSE和内积损失的表现相似;更深的快速权重网络和归一化往往会损害性能,因为它们会导致激活值过大,而残差连接和门控则几乎没有可测量的益处。基于这些发现,我们在1000亿个token上训练了4.1亿参数和14.5亿参数的最佳变体模型,观察到其训练损失和基准性能可与Gated DeltaNet相媲美。

英文摘要

Test-time training (TTT) views sequence modeling as an online learning problem in which fast weights are updated by an internal learning rule. Despite the growing number of TTT variants, existing approaches typically hard-code each variant separately, which makes it difficult to design new TTT methods and to isolate the role of each component. To address this, we propose Modular TTT, a framework that represents the inner learner as a directed acyclic graph and exposes the fast-weight network, loss function, learning rate, weight decay, and normalization as explicit design dimensions. Modular TTT automatically composes primitive-level train-view forward, train-view backward, and causal query-view rules into the full graph-level TTT computation, including the fast-weight state transition. Using Modular TTT, we systematically ablate the components of TTT and find that small learning-rate initialization, weight decay, and a single-layer nonlinearity improve performance, while MSE and inner-product losses perform similarly. Deeper fast-weight networks and normalization tend to hurt performance because they induce excessively large activations, while residual connections and gating provide little measurable benefit. Guided by these findings, we train the best resulting variant as 410M- and 1.45B-parameter models on 100B tokens, and observe training loss and benchmark performance comparable to Gated DeltaNet.

发表机构

  • Shanghai Jiao Tong University(上海交通大学)
  • Shanghai Innovation Institute(上海创新研究院)
  • ByteDance Seed(字节跳动Seed)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑