AI 中文总结
针对测试时语言模型记忆易被后续更新覆盖的问题,提出延迟监督方法,在长间隔后提问并监督答案,在LaCT和DeltaNet上显著提升BABILong与RULER性能。
AI 中文摘要
测试时语言模型在处理输入序列的同时调整一个紧凑的记忆。这一视角涵盖了LaCT中的非线性快速权重学习、DeltaNet中的关联delta规则更新,以及RWKV-7中的广义delta规则状态更新。训练这些模型以预测下一个词元,并不明确要求某个事实在多次后续记忆更新后仍保持可访问。我们研究这种测试时记忆的延迟监督:在后训练期间,仅在经过一段较长的无关事件间隔后,才提出一个基于模拟器的问题,并对其答案与普通的下一词元预测一同进行监督。问题在一次性分支上进行评估,因此其答案永远不会进入持续的事件流。这种构造区分了保留与修订:保留的事实必须在整个延迟期间保持有效,而修订的事实则必须以其最新值来回答。我们在LaCT-760M和普通DeltaNet-1.3B上,使用TextWorld训练轨迹以及共享的BABILong和RULER评估面板来评估这种方法,并包含一个单独报告的RWKV-7对比。相对于仅事件训练,延迟问答使LaCT在BABILong上提高了5.48个百分点,DeltaNet提高了1.32个百分点;在单针RULER上分别提高了1.45和3.27个百分点。RWKV-7对比在其自身面板上报告了4.60和7.00个百分点的提升。这些结果支持将延迟语义监督作为可用测试时记忆的一种实用外部训练目标,但尚不清楚收益中有多少具体来自延迟,而非一般的问答和答案终止监督。
英文摘要
Test-time language models adapt a compact memory while processing the input sequence. This perspective encompasses nonlinear fast-weight learning in LaCT, associative delta-rule updates in DeltaNet, and generalized delta-rule state updates in RWKV-7. Training these models to predict the next token does not explicitly require a fact to remain accessible after many subsequent memory updates. We study delayed supervision for this test-time memory: during post-training, ask a simulator-grounded question only after a long interval of unrelated events, and supervise its answer alongside ordinary next-token prediction. Questions are evaluated on disposable branches, so their answers never enter the continuing event stream. The construction distinguishes retention from revision: a retained fact must remain valid throughout the delay, whereas a revised fact must be answered with its latest value. We evaluate this approach on LaCT-760M and plain DeltaNet-1.3B using TextWorld training trajectories and shared BABILong and RULER evaluation panels, and include a separately reported RWKV-7 comparison. Relative to event-only training, delayed QA improves BABILong by 5.48 percentage points for LaCT and 1.32 points for DeltaNet, and single-needle RULER by 1.45 and 3.27 points, respectively. The RWKV-7 comparison reports gains of 4.60 and 7.00 points on its own panels. These results support delayed semantic supervision as a practical outer training objective for usable test-time memory, while leaving open how much of the benefit derives specifically from delay rather than general question-answering and answer-termination supervision.
Comments11 pages, 3 figures. Equal contribution by both authors. Work done at MIT CSAIL