GradLev:通过协态预测实现令牌并行的测试时训练
GradLev: Token-Parallel Test-Time Training Via Costate Prediction
查看机构详情
- The University of Texas at Austin(德克萨斯大学奥斯汀分校)
机构由 AI 辅助整理,请以论文原文为准。
浏览论文内容
中文总结 AI 辅助
GradLev通过协态预测实现令牌并行的测试时训练,利用在线梯度下降的精确并行扫描特性,以一致性损失监督预测器,保证精确恢复顺序在线学习者,并在部署时丢弃辅助网络。
中文摘要 AI 辅助
测试时训练(TTT)允许模型在推理时通过在每个观察到的令牌后更新权重来改进其预测。然而,顺序梯度写入使得并行训练变得困难。我们观察到,给定层输入和激活梯度(协态),在线梯度下降对前向评估和反向传播均允许精确的并行扫描。GradLev利用了这一对偶性:一个因果辅助网络并行地预测所有令牌的协态;结合扫描计算自适应权重和前向激活,并将梯度反向传播;由此产生的梯度目标通过一致性损失监督预测器。精确一致性保证了顺序在线学习者的精确恢复。在部署时,辅助预测器被丢弃,模型通过逐令牌的前向和反向传递原生更新。
英文摘要
Test-time training (TTT) allows a model to improve its predictions at inference time by updating weights after every observed token. However, sequential gra- dient writes make parallel training difficult. We observe that, given layer inputs and activation gradients (costates), online gradient descent admits exact parallel scans for both forward evaluation and reverse backpropagation. GradLev lever- ages this duality: a causal auxiliary network predicts costates across all tokens in parallel; associative scans compute the adapted weights and forward activations and propagate gradients backward; and the resulting gradient targets supervise the predictor via a consistency loss. Exact consistency guarantees exact recovery of the sequential online learner. At deployment, the auxiliary predictor is discarded, and the model updates natively via token-by-token forward and backward passes.