arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

DeltaTTT:非线性循环记忆的逐层优化

DeltaTTT: Layerwise Optimization for Nonlinear Recurrent Memory

Yining Li, Dongchen Han, Jie Fu, Gao Huang

arXiv 2610.08553首次发表:更新:

发表机构

Qiuzhen College, Tsinghua University; LeapLab, Tsinghua University; IQuest Research(清华大学求真书院; 清华大学LeapLab实验室; IQuest研究院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对非线性记忆在序列测试时训练中优化困难的问题,提出DeltaTTT,采用逐层学习与状态相关delta规则,在保留非线性读出的同时支持并行计算,并在DeltaNet和LaCT上提升语言建模与检索性能。

AI 中文摘要

序列测试时训练通过连续更新来调整记忆网络,每次更新基于网络先前状态计算内循环梯度。直觉上,这种状态依赖性应使每次更新能够考虑记忆已学到的内容,并更好地整合新信息。然而,我们发现这一预期优势在非线性记忆中并未持续显现:固定基线的并行TTT基线优于其串行对应版本。我们的探索性实验指出了关键潜在困难:在单次序列遍历中,非线性记忆可能比线性记忆更难优化。为缓解这一优化困难,我们提出了DeltaTTT,将两层记忆网络的联合内循环优化替换为逐层学习。每层被分配一个局部预测目标,并通过状态相关的delta规则进行更新。该公式保留了非线性读出,同时支持分块并行计算。在DeltaNet和LaCT骨干上的实验表明,在语言建模和检索任务上相比其循环基线有所改进。

英文摘要

Sequential test-time training adapts a memory network through successive updates, each computing an inner-loop gradient based on the network's previous state. Intuitively, this state dependence should allow each update to account for what the memory has already learned and better incorporate new information. However, we find that this expected advantage does not consistently materialize in nonlinear memories: a fixed-base parallel TTT baseline outperforms its serial counterpart. Our exploratory experiments point to a key underlying difficulty: nonlinear memories can be harder to optimize than linear ones within a single pass over the sequence. To alleviate this optimization difficulty, we introduce DeltaTTT, which replaces joint inner-loop optimization of a two-layer memory network with layerwise learning. Each layer is assigned a local prediction target and updated through a state-dependent delta rule. This formulation retains a nonlinear readout while enabling chunkwise parallel computation. Experiments on DeltaNet and LaCT backbones show improvements in language modeling and retrieval over their recurrent baselines.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑