arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.12218cs.CLcs.AI

信息丰度悖论:长上下文训练会破坏参数化知识

Information Abundance Paradox: Long-Context Training Undermines Parametric Knowledge

发表机构约翰斯·霍普金斯大学
查看机构详情
  • Johns Hopkins University(约翰斯·霍普金斯大学)

机构由 AI 辅助整理,请以论文原文为准。

Arda Uzunoglu, Benjamin Van Durme, Daniel Khashabi

首次发表
浏览论文内容

中文总结 AI 辅助

该研究提出信息丰度悖论,发现长上下文训练会使模型减少参数化知识编码、增加对上下文的依赖,预训练和微调中均存在性能先升后降或鲁棒性降低的现象,表明近无穷大上下文缩放并非仅靠更多数据。

中文摘要 AI 辅助

大型语言模型正日益以覆盖文档、代码仓库和交互历史的长上下文进行训练和部署,这种缩放反映了一种隐含假设:在更长的上下文上训练只会通过为模型提供更丰富的证据来帮助它。我们通过研究上下文窗口如何塑造模型的学习模式,在参数内化和上下文依赖之间切换,对这一观点提出挑战。我们提出信息丰度悖论,该假说认为训练上下文中的大量相关信息会降低将该信息编码为参数的动机,从而增加对上下文的依赖。在长文档的预训练中,增大上下文窗口可提升语言建模、自然语言理解和闭卷多项选择问答(MCQA)的性能,直至达到中间最优值,此后性能会持续下降。在监督微调中,更多与任务相关的训练时上下文在有支持性上下文时会提升性能,但在测试时上下文缺失或具有误导性时会降低鲁棒性。我们的分析表明,这种行为出现的原因是更长的上下文提供了复杂度更低的解决方案。从机制上看,使用信息丰富的上下文进行训练会将梯度压力从通常与参数化知识相关的前馈网络转移到注意力模块,而因果干预表明,这种转移会增加推理时对上下文的依赖。总体而言,这些发现支持信息丰度悖论,并表明向近无穷大上下文缩放并非仅仅是提供更多数据的问题,即便高质量长上下文数据十分丰富亦是如此。

英文摘要

Large language models are increasingly trained and deployed with long contexts that span documents, code repositories, and interaction histories. This scaling reflects the implicit assumption that training on longer contexts will only help the model by exposing it to richer evidence. We challenge this view by studying how the context window shapes a model's mode of learning, shifting it between parametric internalization and contextualization. We propose the Information Abundance Paradox, which hypothesizes that abundant relevant information in the training context can reduce the incentive to encode that information parametrically, thereby increasing reliance on context. In pretraining with long documents, increasing the context window improves language modeling, natural language understanding, and closed-book MCQA only up to an intermediate optimum, after which performance consistently declines. In supervised fine-tuning, more task-relevant train-time context improves performance with supporting context, but reduces robustness when context is absent or misleading at test time. Our analysis suggests that this behavior arises when longer context provides a lower complexity solution. Mechanistically, training with informative context shifts gradient pressure from feed-forward networks, often linked to parametric knowledge, toward attention modules, and causal interventions show that this shift increases reliance on context during inference. Overall, these findings support the Information Abundance Paradox and suggest that scaling toward near-infinite context is not simply a matter of supplying more data, even when high-quality long-context data is abundant.

↑