发表机构
Estonian Entrepreneurship University of Applied Sciences(爱沙尼亚创业应用科学大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
通过重建Laurito等人的5×5生成器-选择器矩阵并拟合固定效应模型,发现模型对模型生成文本有共同偏好,但未检测到显著的自模型溢价,即模型不偏爱自己的输出。
AI 中文摘要
Laurito 等人(PNAS 2025)的研究表明,当大型语言模型在同一个产品、论文或电影的两段描述之间进行选择时,它们更倾向于选择由语言模型撰写的描述,而非由人类撰写的描述,其偏好程度远超人类评审员的判断。他们的实验设计将五个生成器与作为选择器的相同五个模型交叉配对,这引出了一个该论文未重点提及的第二个问题:选择器是否会偏好来自其自身模型的文本,而这种偏好超出了生成器和选择器主效应所能预测的范围?我们从作者公开存储库中的逐项目计数重建了三个 5×5 矩阵(共 21,828 次有效试验;每个单元格都与已发表的值完全匹配),并拟合了一个包含自模型项 gamma 的双向固定效应模型,通过选择器的 120 种重新标记的精确置换检验进行测试。自模型溢价在产品上为 +0.013(精确单侧 p = 0.24),在论文摘要上为 -0.010(p = 0.74),在电影上为 +0.054(p = 0.07),合并后为 +0.019(p = 0.14;95% 置信区间为 -0.008 至 0.046)。GPT-3.5 和 GPT-4 这一配对在三个数据集中的同供应商项均为负值。位置偏差会使单个单元格在任一方向上最多移动 0.42 个百分点,而在移除顺序驱动的项目后,自模型对比保持不变。该设计能够以 82%(产品)、88%(论文)、42%(电影)和 97%(合并)的统计功效检测到 0.05 的溢价;在 80% 功效下可检测的最小效应为合并后 0.034。这种缺失在约 0.04 个百分点以下具有信息量,而在更低水平则无法提供信息。Tan 等人(ACL 2024)的 4×4 矩阵在 24 种重新标记中给出了 gamma = +0.148,其 p 值达到最小,同族项的大小相同。Laurito 等人的主要结果依然成立:模型对模型生成的文本有共同偏好,在产品上,每个选择器选择 GPT-4 描述的比例为 77% 至 95%。这些数据未显示的是,模型能够识别并偏爱自己的文本。
英文摘要
Laurito et al. (PNAS 2025) showed that large language models choosing between two descriptions of the same product, paper or film prefer the description written by a language model over the one written by a person, by a wide margin over what human judges do. Their design crosses five generators with the same five models as selectors, which permits a second question the paper does not headline: does a selector prefer text from its own model beyond what the generator and selector main effects predict? We rebuild the three 5x5 matrices from the per-item counts in the authors' public repository (21,828 valid trials; every cell matches the published value) and fit a two-way fixed-effects model with an own-model term gamma, tested by the exact permutation test over the 120 relabellings of the selectors. The premium is +0.013 on products (exact one-sided p = 0.24), -0.010 on paper abstracts (p = 0.74), +0.054 on films (p = 0.07) and +0.019 pooled (p = 0.14; 95% interval -0.008 to 0.046). The same-vendor term for the GPT-3.5 and GPT-4 pair is negative in all three datasets. Position bias moves single cells by up to 0.42 share points in either direction, and the own-model contrast is unchanged once order-driven items are removed. The design would have detected a premium of 0.05 with 82% (products), 88% (papers), 42% (films) and 97% (pooled) power; the minimum detectable effect at 80% power is 0.034 pooled. The absence is informative down to about 0.04 share points and silent below that. The 4x4 matrix of Tan et al. (ACL 2024) gives gamma = +0.148 at the smallest p its 24 relabellings allow, with a same-family term of the same size. The main result of Laurito et al. stands: models share a taste for model-written text, with GPT-4's descriptions chosen 77% to 95% of the time by every selector on products. What these data do not show is a model recognising and favouring its own prose.
Comments10 pages, 4 figures, 3 tables. Reanalysis of publicly available generator-by-selector matrices