arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

学术论文推荐用对话式搜索引擎中的权威偏差

Authority Bias in Conversational Search Engines for Academic Paper Recommendation

Uthman Jinadu, Parsa Ghazvinian, Anjila Budathoki, Benjamin M. Ampel, Rajshekhar Sunderraman, Yi Ding

arXiv 2609.00248首次发表:更新:

发表机构

Georgia State University; University of Tennessee, Knoxville(佐治亚州立大学; 田纳西大学诺克斯维尔分校)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究针对学术论文推荐的对话式 LLMs,发现其存在基于作者声望等的权威偏差,不同模型偏差差异大,提示级去偏仅部分有效,且存在言行差距致表面审计低估偏差。

AI 中文摘要

大型语言模型(LLMs)正越来越多地被用作学术文献的对话式搜索引擎,但尚未对它们是依据内容还是权威信号来评判论文进行因果测试。我们研究权威偏差:即系统地依据作者声望、 venues( venues 保留原词,可理解为学术会议/期刊)和引用量而非论文内容来偏好论文的现象。在保持标题和摘要不变的情况下,我们针对八种 LLMs(五种开放权重模型和三种前沿封闭权重模型),在上下文式单轮 top-1 推荐设置中,设置三种反事实条件(原始、翻转、增强)来改变权威元数据。实验表明,权威偏差是显著且有方向的,不同模型间差异明显,且仅能通过提示级去偏部分解决。我们还发现了言行差距:去偏指令抑制权威提及的速度远快于权威驱动的翻转,因此表面审计会系统性低估行为偏差。

英文摘要

Large Language Models (LLMs) are increasingly used as conversational search engines for academic literature, yet whether they judge papers on content or on authority signals has not been tested causally. We investigate authority bias: systematic preference for papers based on author prestige, venue, and citations rather than content. Holding title and abstract constant, we vary authority metadata across three counterfactual conditions (original, flipped, boosted) over eight LLMs (five open-weight and three frontier closed-weight) in an in-context, single-turn, top-1 recommendation setting. Our experiments show that authority bias is substantial and directional, varies markedly across models, and is only partially addressable through prompt-level debiasing. We further document a say-do gap: debiasing instructions suppress authority mentions far faster than authority-driven flips, so surface auditing systematically underestimates behavioral bias.

CommentsAccepted at EMNLP 2026 Main Conference

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑