arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

下限、上限与融合差距:机器能预测多少群体阅读注意力?

Floor, Ceiling, and the Fusion Gap: How Much of Crowd Reading Attention Can Machines Predict?

Kazuki Nakayashiki, Keisuke Watanabe

arXiv 2608.01704首次发表:更新:

发表机构

Glasp Inc.(Glasp公司)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究构建了群体阅读注意力预测任务的上下限与融合差距,发现前沿模型仅能捕捉部分差距,多模型融合可提升性能,蒸馏后的8B学生模型能保留融合的大部分优势,证实人群信号存在于文档级结构中。

AI 中文摘要

基准分数若不了解简单方法能达到的结果和最优方法能达到的结果,便毫无意义。我们针对一项任务构建了上下限,该任务采用了一种罕见的真值类型:预测由120份网页文档组成的内容中,一群读者(出于自身目的高亮、无报酬、无指令且互不了解)标记的句子。下限是朴素截断法(取前导句);上限是半分神谕:用一半人群预测另一半人群。二者间的差距为平均精度(AP)+0.2028(95%置信区间[+0.1698, +0.2342],按领域聚类),三项发现构成该差距的结构:其一,差距具有语义性,位置和长度特征仅能恢复其中5%;其二,前沿语言模型零样本可达到该差距的35%-53%,远高于经典基线但远低于人群;最先进的提示压缩器LLMLingua-2表现低于下限,与随机选择无差异;其三,五个前沿排名与位置先验的未加权跨供应商融合可达到60%,比最优单模型高出+0.0159(95%置信区间[+0.0044, +0.0269],Holm检验p=0.019),该增益在剔除最优成员、半分臂选择、提示改写及标签、门控、种子扰动后仍存在,且经217份独立文档的预注册复现得到证实(+0.0179,Holm检验p=0.042)。最后,该 bracket 可压缩:将融合结果蒸馏为一个读取完整文档的开放权重8B学生模型,可保留融合90%的优势,与最强单前沿模型达到统计均等(+0.0070,95%置信区间[-0.0068, +0.0200]);而仅读取局部上下文的学生模型仅保留63%的优势——人群的信号存在于文档级结构中,目前已知的最低成本改进方法是询问多个不同模型并取平均。

英文摘要

A benchmark score means nothing without knowing what a trivial method achieves and what the best possible method could achieve. We construct both bounds for a task with a rare kind of ground truth: predicting which sentences a crowd of readers -- highlighting for their own purposes, unpaid, uninstructed, and blind to each other -- marked in 120 web documents. The floor is naive truncation (lead); the ceiling is a split-half oracle: half the crowd predicting the other half. The gap between them is +0.2028 AP [+0.1698, +0.2342, domain-clustered], and three findings structure it. First, the gap is semantic: position and length features recover 5% of it. Second, frontier language models reach 35-53% of it zero-shot -- far above classical baselines, far below the crowd; a state-of-the-art prompt compressor (LLMLingua-2) lands below the floor, indistinguishable from random selection. Third, an unweighted cross-vendor fusion of five frontier rankings plus a position prior reaches 60%, beating the best single model by +0.0159 [+0.0044, +0.0269; Holm p=0.019] -- a gain that survives ablation of its best member, split-half arm selection, prompt paraphrase, and label, gate, and seed perturbations, and was CONFIRMED by a pre-registered replication on 217 independent documents (+0.0179, Holm p=0.042). Finally, the bracket compresses: distilling the fusion into one open-weight 8B student that reads the whole document retains 90% of the fusion's edge and reaches statistical parity with the strongest single frontier model (+0.0070 [-0.0068, +0.0200]), where a local-context student retains only 63% -- the crowd's signal lives in document-level structure, and the cheapest known improvement is to ask several different models and average.

Comments8 pages. Ancillary files include the pre-registrations, hostile-audit records, verification scripts, and the aggregate artifacts every reported number is generated from

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑