学习该记住什么:基于上下文蒸馏的测试时训练
Learning What to Remember: Test-Time Training via Context Distillation
浏览论文内容
中文总结 AI 辅助
该研究针对现有测试时训练未考虑保留信息未来效用的问题,提出TTCD框架及IP-TTCD变体,实验显示其在长上下文语言建模任务中性能优于多个基准方法,可助力架构持续学习。
中文摘要 AI 辅助
有效的长上下文建模不仅仅是保留更多过去的信息,还在于保存后续可能相关的信息。测试时训练(Test-Time Training, TTT)是一种适用于长上下文建模的在线参数更新方法,但现有TTT方法仅优化重构或在线适应目标,未考虑保留信息的未来效用。本文提出测试时上下文蒸馏(Test-Time Context Distillation, TTCD)框架,引入自监督目标以分配有限的内存容量供未来使用。具体而言,TTCD利用长窗口教师模型监督短窗口学生模型的快速权重,二者的隐藏状态差异提供了密集的自监督信号,引导模型记忆对未来token预测至关重要的上下文信息。本文聚焦原位变体:原位TTCD(In-Place TTCD, IP-TTCD),该方法将现有MLP参数作为快速权重。在长上下文语言建模任务上的实验表明,当从头预训练时,IP-TTCD的性能始终优于DeltaNet、门控DeltaNet、滑动窗口注意力和TTCD。此外,IP-TTCD允许预训练的Transformer模型在推理时通过持续预训练调整参数,仅需轻量级架构增强即可获得长上下文能力,本文的结果表明TTCD是迈向架构持续学习的一步。
英文摘要
Effective long-context modeling is not merely about retaining more of the past, but about preserving the information that may prove relevant later. Test-time training (TTT) is an appealing approach that performs online parameter updates for long-context modeling, yet existing TTT methods only optimize either reconstruction or online adaptation objectives without considering the future utility of retained information. In this work, we propose \textbf{T}est-\textbf{T}ime \textbf{C}ontext \textbf{D}istillation (TTCD), a TTT framework that introduces a self-supervised objective for allocating limited memory capacity for future use. Specifically, TTCD uses a long-window teacher to supervise the fast weights of a short-window student, where the hidden-state discrepancy between them offers a dense, self-supervised signal guiding the model to memorize the contextual information crucial for future token predictions. We focus on an in-place variant: In-Place TTCD (IP-TTCD), which uses the existing MLP parameters as the fast weights. Experiments on long-context language modeling tasks show IP-TTCD consistently outperforms DeltaNet, Gated DeltaNet, sliding-window attention, and TTT when pre-trained from scratch. Furthermore, IP-TTCD allows pre-trained transformer models to adapt their parameters during inference through continual pre-training, gaining long-context capabilities with only a lightweight architectural augmentation. Our results position TTCD as a step toward architectural continual learning.
发表机构
- Princeton University(普林斯顿大学)
- UC Berkeley(加州大学伯克利分校)
- UC Santa Cruz(加州大学圣克鲁兹分校)
- Carnegie Mellon University(卡内基梅隆大学)
机构由 AI 辅助整理,请以论文原文为准。