发表机构
King Abdullah University of Science and Technology (KAUST)(阿卜杜拉国王科技大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究通过因果分解揭示测试时训练中自生成反馈会破坏长时域适应,提出固定生成等方法消除损害,并强调在保留更新前需在独立证据上验证预测。
AI 中文摘要
测试时训练(TTT)允许模型在推理期间将信息存储在其权重中。然而,当模型从自身输出中学习时,每次更新也会改变生成下一个训练示例的模型。在128K token的流中,使用三种TTT-E2E模型配置(标记为125M、760M和3B)时,保留生成的文本更新会恶化对独立人类书写文本的预测。当Adam更新Qwen3-4B的现有权重时,也会出现同样的失败。相同的更新机制可以改进真实文本,因此写作本身并非失败原因。三个匹配的比较追踪了因果路径。固定生成通过使用冻结模型生成训练块,在125M和760M下消除了超过98%的损害。记录重放将读取退化文本造成的损失与在其上更新所存储的额外损失分开。配对的一次更新比较随后显示了局部冲突:更新能更好地预测其来源,但对新真实文本的预测更差。这种代价在闭环适应后增长,少数轨迹造成了大部分重大失败。最后,结算在提交前评估候选状态在独立真实文本上的表现。它在125M和760M下留下了0.07和-0.02 nats的平均终点差距,同时保留了真实文本适应。这些结果激励在保留更新前检查独立证据上的预测。
英文摘要
Test-time training (TTT) lets a model store information in its weights during inference. When the model learns from its own output, however, each update also changes the model that generates the next training example. Across 128K-token streams, retaining generated-text updates worsens prediction on independent human-written text with three TTT-E2E model configurations (labeled 125M, 760M, and 3B). The same failure occurs when Adam updates Qwen3-4B's existing weights. The same update mechanisms can improve on real text, so writing itself is not the failure. Three matched comparisons trace the causal pathway. Fixed Generation removes over 98% of the damage at 125M and 760M by using a frozen model to generate training chunks. Recorded Replay separates the loss caused by reading degraded text from the additional loss stored by updating on it. A paired one-update comparison then shows the local conflict: an update predicts its source better but new real text worse. This cost grows after Closed Loop adaptation, with a few trajectories accounting for most large failures. Finally, Settlement evaluates the candidate state on independent real text before commitment. It leaves mean endpoint gaps of 0.07 and -0.02 nats at 125M and 760M while retaining real-text adaptation. These results motivate checking prediction on independent evidence before retaining an update.