发表机构
Glasp Inc.(Glasp公司)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究针对上下文压缩评估的混淆问题,提出匹配位置长度的方法,发现语言模型重要性排名对齐度优于人类读者外的基准,且前期某结论无法在当前语料库复现。
AI 中文摘要
上下文压缩会在语言模型读取文档前丢弃大部分内容,通常通过下游任务准确率评估,这使得另一模型成为判断内容重要性的依据。自然的社会标注提供了非循环参考:多人独立标记同一页面的段落。但直观指标——压缩器保留的群体标记句子比例,存在双重混淆:群体标记集中在文档前部,且被标记句子更长,因此任何偏向早出现或长句子的方法都会表现良好,与读者真实判断无关。我们通过将每个被标记句子与同一文档中相同相对深度、相同文档内长度排名的未标记句子匹配,消除上述两个混淆因素;并基于仅由位置和长度构成的合成零假设校准每个估计量,这一步至关重要,因为仅按深度分层会在无效应的零假设中产生20%-36%的假阳性。在120份网页文档(每份至少12名独立读者)上,语言模型重要性排名保留了38.4%的群体标记句子,而其匹配的未标记邻居仅保留19.9%,富集度为+0.196(95%置信区间[+0.148, +0.239]),在不假设聚类的精确随机化检验下p值为0.0005,且跨供应商可重复。基于位置的朴素截断法富集度仅为+0.003。为提供规模参照:在相同预算下,针对排除该读者后重新计算的群体标签,单个人类读者的表现为+0.182,与GPT-5.4(+0.002,95%置信区间[-0.081, +0.088])无差异,低于Claude Opus 5。经典方法并非无效:Luhn1958年的启发式方法富集度为+0.088,说明读者选择的部分信息可通过计数单词恢复;额外加入词汇中心性条件仅减少0.010的富集度,因此该一致性与词汇中心性无关。我们还报告,本团队前期研究的一项结论在该语料库上无法复现。
英文摘要
Context compression discards most of a document before a language model reads it, and is normally evaluated by downstream task accuracy - which makes another model the judge of what mattered. Naturalistic social highlighting offers a non-circular reference: many people independently marking passages on the same page. But the obvious metric, the fraction of crowd-marked sentences a compressor keeps, is confounded twice: crowd marks are front-loaded and crowd-marked sentences are longer, so any method favouring early or long sentences scores well regardless of readers. We remove both by matching each marked sentence against unmarked sentences of the same document at equal relative depth and equal within-document length rank, and we calibrate every estimator on synthetic nulls built from position and length alone - a step that matters, since depth-only stratification returns a false positive on 20-36% of nulls containing no effect. On 120 web documents (at least 12 independent readers each), a language-model importance ranking keeps 38.4% of crowd-marked sentences against 19.9% of their matched neighbours: an enrichment of +0.196 [+0.148, +0.239], at p = 0.0005 under an exact randomization test that assumes nothing about clustering, and replicated cross-vendor. Naive truncation, whose keep rule is position, correctly falls to +0.003. To give the number a scale: scored identically, on the same budget, against a crowd label recomputed to exclude them, a single human reader reaches +0.182 - indistinguishable from GPT-5.4 (+0.002 [-0.081, +0.088]) and below Claude Opus 5. Classical methods are not null - Luhn's 1958 heuristic reaches +0.088 - so reader selection is partly recoverable by counting words; conditioning additionally on lexical centrality removes only 0.010, so the agreement is not centrality. We also report that a claim in our own prior work does not reproduce on this corpus.
Comments15 pages, 7 tables. Analysis code and de-identified artifacts included as ancillary files; five of six scripts reproduce the paper's numbers from the shipped artifacts alone. Reports claims from our own prior work that this corpus does not reproduce, and lists twelve claims withdrawn during internal adversarial review in Appendix A