arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

语言模型彼此意见一致,但与读者不一致

Language Models Agree With Each Other, Not With Readers

Kazuki Nakayashiki, Keisuke Watanabe

arXiv 2607.29274首次发表:更新:

发表机构

Glasp Inc.(Glasp公司)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究以非研究专用的读者标记为参考,发现不同语言模型间的一致性高于模型与读者间的一致性,且模型一致性与模型规模等因素相关,样本外测试也验证了该结论。

AI 中文摘要

声称语言模型同质化的研究通常以该研究收集的人类判断为衡量标准,这使得人类一方成为设计的产物:获得模型指令的众包工作者正在运行模型的提示。我们针对一个并非为此目的构建的人类参考来衡量收敛性——120篇网络文档的2523组读者标记,这些标记是人们在一个平台上出于自身原因高亮显示的,该平台默认关闭他人标记的叠加功能。一致性是两个大小匹配的句子集之间的重叠,减去每个句子在自身深度和长度范围内重新采样时预期的重叠。零假设的校准是通过演示而非断言:涉及随机基线的每一对都落在0.006以内。在中位数文档中,每一方都标记了70个句子中的14个;两名读者共享4.1个,两个模型共享8.7个。在涵盖11个供应商、3个国家和两种权重机制的18个模型分支中,153个模型对的中位数相对于人类基准为+0.093,而人类基准的中位数为+0.040,其中99个完全高于人类区间。来自竞争实验室的两个前沿模型达到+0.203,是GPT-4o在第二次调用时与自身达成一致的两倍。这种效应并非由确定性、提示措辞、程序、供应商或路由导致,且是分级的:最小的模型与人类水平一致。没有模型与读者的一致性明显多于读者之间的一致性,且在相同深度和长度下,没有表面特征能区分它们的选择。倍数取决于程序,而顺序则不然:模型被裁剪到最清晰的集合,而读者的标记是其标记内容的随机抽取,同等程度地削弱模型会使差距减半但无法消除。在本分析之后发布的四个模型上进行样本外测试,针对预先确定的预测,没有一个能通过人类区间。由多个模型模拟的总体并非多个总体。

英文摘要

Claims that language models homogenise are usually measured against human judgements collected for the study, which makes the human side an artifact of the design: a crowdworker given the model's instruction is running the model's prompt. We measure convergence against a human reference nobody built for the purpose -- 2,523 reader mark sets across 120 web documents, produced by people highlighting for their own reasons on a platform where the overlay of others' marks is off by default. Agreement is the overlap between two size-matched sentence sets minus the overlap expected when each is resampled within its own depth-and-length bands. The null's calibration is demonstrated, not asserted: every pair involving a random baseline lands within 0.006 of zero. On the median document each party names 14 sentences of 70; two readers share 4.1 and two models 8.7. Across 18 model arms spanning 11 vendors, 3 countries and both weight regimes, the median of 153 model pairs is +0.093 against a human yardstick of +0.040, and 99 sit entirely above the human interval. Two frontier models from rival labs reach +0.203, twice what GPT-4o agrees with itself on a second call. The effect is not determinism, prompt wording, procedure, vendor or routing, and it is graded: the smallest models agree at the human level. No model agrees with readers detectably more than a reader does, and at equal depth and length no surface feature separates their choices. The multiples are procedure-dependent and the ordering is not: models are cut to their sharpest set while a reader's is a random draw from what they marked, and blunting the models alike halves the gap without closing it. Tested out of sample on four models released after this analysis, against predictions fixed beforehand, none clears the human interval. A population simulated from several models is not several populations.

Comments18 pages. Ancillary files include all three pre-registrations, every analysis script and every result artifact; the paper contains no numeric literal for a measured value and make-numbers.py regenerates all of them from the shipped artifacts alone

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑