arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

当AI撰写时,谁会被引用?语言模型间的引用单一文化证据

When AI Writes, Who Gets Cited? Evidence of Citation Monoculture Across Language Models

Sina Alemohammad, Denghui Zhang, Bolong Tang, Anthony Qin, Gengchen Mai, Ahmed Abbasi, Richard Baraniuk, Zhangyang Wang

arXiv 2608.19230首次发表:更新:

发表机构

The University of Texas at Austin; Stevens Institute of Technology; Washington University in St. Louis; University of Notre Dame; Rice University(德克萨斯大学奥斯汀分校; 史蒂文斯理工学院; 圣路易斯华盛顿大学; 圣母大学; 莱斯大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究发现,即使引用真实,当前语言模型仍会产生引用单一文化,需改变其共享偏好图而非仅均衡检索或混合厂商。

AI 中文摘要

随着语言模型从起草散文转向使用工具调用运行文献搜索智能体,伪造的参考文献变得更容易被发现和约束。更棘手的问题出现在所有候选文献均为真实的情况下:不同模型仍可能选择相同的狭窄子集,产生引用单一文化,且无任何单篇引用错误。我们在120篇真实论文上分离出该效应。来自三家厂商的11个模型,从包含30篇论文的均匀随机面板中选择至多10篇,这些论文拥有真实标题和摘要,但作者为伪造、年份被重新分配、 venue( venues)和引用数被隐藏。每次运行都与同一面板及实际预算下的无差别选择进行比较。所有11个模型都呈现出明显的集中趋势:前十分位获得23.3%-30.2%的引用,而随机选择下为15.6%;一个成分解释了其偏好图中68%-73%的变异;跨厂商一致性几乎与厂商内一致性相当。将该任务形式化为固定预算子集选择,我们将这些模式转化为可识别机制:可交换性边界对每个模型都拒绝无偏好图的选择器;谱分解解释了为何最优交叉拟合混合模型仍保留55%的超额;稀有性定理预测了我们在面板内验证的递归竞争效应。受控释义、内容槽交叉及设计重采样将GPT-5 mini约90%的偏好图变异归因于论文内容。8位领域专家在相同盲选面板及相同预算下选择,未表现出可比的共享偏好,而仅选择模式下模型的集中性仍存在。即使每篇参考文献均为真实且每篇论文同等可见,当前语言模型仍对科学注意力施加了共同的内容级筛选。因此,仅均衡检索或混合厂商是不够的,必须改变共享偏好图本身。

英文摘要

As language models move from drafting prose to running literature-search agents with tool calls, fabricated references are becoming easier to catch and constrain. The harder failure begins after every candidate is real: different models may still select the same narrow subset, producing citation monoculture without any single citation being wrong. We isolate this effect on 120 real papers. Eleven models from three vendors choose at most ten papers from uniformly random panels of thirty, with real titles and abstracts but fabricated authors, reassigned years, and hidden venues and citation counts. Each run is compared with indifferent selection on the same panel and realized budget. All eleven models concentrate sharply: the top decile receives 23.3-30.2% of citations against 15.6% under the null, one component explains 68-73% of variation across their preference maps, and cross-vendor agreement nearly matches within-vendor agreement. Formalizing the task as fixed-budget subset selection, we turn these patterns into identifiable mechanisms: an exchangeability bound rejects a mapless selector for every model, a spectral decomposition explains why the best cross-fitted mixture still retains 55% of the excess, and a rarity theorem predicts the recursive competition effect we verify within panels. Controlled paraphrase, content-slot crossover, and design resampling attribute about 90% of GPT-5 mini's map variance to paper content. Eight domain experts selecting from the same blinded panels under the same cap show no comparable shared preference, while model concentration persists in selection-only mode. Even when every reference is real and every paper is equally visible, current language models impose a common content-level filter on scientific attention. Equalizing retrieval or mixing vendors is therefore insufficient; the shared preference map itself must be changed.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑