发表机构
Meta AI; University of Virginia(Meta AI; 弗吉尼亚大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究针对长上下文语言模型处理中单纯扩窗口不能有效利用长输入、准确率易降的问题,提出自引导测试时训练(S-TTT)方法,通过模型识别证据跨度并仅对选定跨度训练,在两个基准上提升了模型准确率。
AI 中文摘要
长上下文处理对大语言模型日益重要,单纯扩展上下文窗口不能保证有效利用长输入。输入长度增加时准确率常下降,模型难以识别和使用与问题最相关的证据。测试时训练(TTT)是提高长上下文利用率的有效方法,但应用于整个长上下文成本过高,在随机采样跨度上进行适应会引入严重噪声。我们的初步研究表明TTT对训练跨度质量高度敏感。基于此,我们提出自引导TTT(S-TTT)方法,在适应前模型识别应学习的证据跨度,并仅对选定跨度应用标准语言建模训练目标。在两个具有挑战性的长上下文推理基准上,S-TTT提高了Qwen3-4B-Thinking-2507和Llama-3.1-8B-Instruct的准确率,相对提高高达15%。
英文摘要
Long-context processing has become increasingly important for large language models (LLMs), but simply extending the context window does not guarantee effective utilization of long inputs. As input length grows, accuracy often degrades, indicating that models still struggle to identify and use the evidence most relevant to a question. A promising way to improve long-context utilization is test-time training (TTT), which treats the test context as a training example for instance-specific parameter adaptation. However, applying TTT to the entire long context is prohibitively expensive, while adapting on randomly sampled spans introduces severe noise. Because most spans in a long context are irrelevant to the specific question, training on them may even degrade the base model's performance. Our preliminary study shows that TTT is highly sensitive to training-span quality: on LongBench-v2, TTT on randomly sampled spans hurts performance, whereas TTT on oracle spans substantially improves it. Motivated by this, we propose a simple method, Self-Guided TTT (S-TTT): before adaptation, the model identifies the evidence spans it should learn from, and the standard language-modeling training objective is applied only to those selected spans. On two challenging long-context reasoning benchmarks, LongBench-v2 and LongBench-Pro, S-TTT improves accuracy for both Qwen3-4B-Thinking-2507 and Llama-3.1-8B-Instruct, achieving up to a 15% relative improvement.