干扰项感知截断:在长上下文大语言模型基准测试中分离上下文长度效应与信号损失
Distractor-Aware Truncation: Disentangling Context-Length Effects from Signal Loss in Long-Context LLM Benchmarks
浏览论文内容
中文总结 AI 辅助
本研究通过对比两种截断协议测试长上下文基准,发现朴素截断的性能下降源于信号丢失,干扰项感知截断可保留性能,强调需明确区分信号与干扰项以研究上下文长度效应。
中文摘要 AI 辅助
检索增强和记忆增强语言模型领域的标准观点是:当相关信息得以保留时,更短的上下文表现更好。我们通过在四种上下文保留比例(100%、75%、50%、25%)下,采用两种截断协议对两个长上下文基准测试——BABILong和GraphWalks(BFS)的每个样本进行测试,以验证该观点。第一种是先前许多工作隐含使用的朴素协议:删除提示中间的内容;第二种是干扰项感知协议:识别每个样本的任务相关内容,仅删除其余部分。我们评估了Claude家族的三种规模模型(Haiku 4.5、Sonnet 4.6、Opus 4.7),并测试跨提供商通用性,对来自不同提供商的GPT-5.5应用相同协议,还扩展到另外两个基准测试(MRCR v2、Oolong)。在朴素截断下,分数单调下降(配对Wilcoxon检验,Holm校正后p_adj在所有8个BABILong和GraphWalks单元中均小于0.05);在干扰项感知协议下(该协议从结构上保留信号),性能得以保留或提升:两个较小的Claude模型在BABILong上表现出统计显著提升,而较大模型(Opus 4.7和GPT-5.5)处于全上下文上限。朴素下降及其干扰项感知恢复在GPT-5.5上也得到重复,排除了单一提供商的人为因素。机制很直接:在25%保留比例下,朴素协议中携带答案的内容在不到1%的样本中保留;而在干扰项感知协议中,该内容从结构上得到保留。因此,朴素协议并非对上下文窗口效应的测量,而是对中间删除恰好保留答案的频率的测量。我们得出结论:未来对上下文长度效应的研究必须明确如何区分信号与干扰项,否则最多只能在两个相反假设间产生歧义。
英文摘要
A standard claim in the literature on retrieval-augmented and memory-augmented language models is that shorter context is better when the relevant information is preserved. We test this claim by running every sample of two long-context benchmarks -- BABILong and GraphWalks (BFS) -- at four context-retention fractions (100%, 75%, 50%, 25%) under two truncation protocols. The first is the naive protocol implicitly used in much prior work: drop content from the middle of the prompt. The second is distractor-aware: identify the task-relevant content for each sample and drop only the rest. We evaluate three sizes of the Claude family (Haiku 4.5, Sonnet 4.6, Opus 4.7) and, to test cross-provider generality, GPT-5.5 from a different provider; we apply the same protocol to two further benchmarks (MRCR v2, Oolong). Under naive truncation, score collapses monotonically (paired Wilcoxon, Holm-corrected p_adj < 0.05 in all eight BABILong and GraphWalks cells). Under the distractor-aware protocol -- which preserves the signal by construction -- performance is preserved or improves: the two smaller Claude models show statistically significant gains on BABILong, while the larger models (Opus 4.7 and GPT-5.5) sit at their full-context ceiling. The naive collapse and its distractor-aware recovery replicate on GPT-5.5, ruling out a single-provider artifact. The mechanism is direct: under the naive protocol the answer-bearing content survives in fewer than 1% of samples at 25% retention; under the distractor-aware protocol it is preserved by construction. The naive protocol is therefore not a measurement of context-window effects; it is a measurement of how often middle-removal happens to spare the answer. We conclude that future studies of context-length effects must specify how they distinguish signal from distractor, or they are at best ambiguous between two opposite hypotheses.