arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

语言模型中的权威偏见:来源遵从与用户认同不可互换

Authority Bias in Language Models: Source Deference and User Agreement Are Not Interchangeable

Abhinav Rajeev Kumar, Paras Chopra

arXiv 2609.37616首次发表:更新:

发表机构

Lossfunk(Lossfunk)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究揭示语言模型对权威来源的遵从远高于对用户的认同,且两者可被独立干预,需分开评估。

AI 中文摘要

语言模型倾向于认同用户所断言的内容,而后期训练越来越针对这种谄媚行为,以使模型根据主张本身的价值来评估,而非遵从用户。然而,当错误答案被归因于经过验证的来源时,同一模型却表现出远为更高的顺从性,而检索结果、工具输出和基于检索的内容通常正是以这种方式呈现信息。我们在五个开放权重模型家族和三个封闭API上衡量了这一差距。在八个模型中的七个里,一条支持错误答案的单一已验证来源注释使基线正确响应中的45-88%发生翻转,且顺从性随注释听起来更具权威性而上升。来源遵从与用户认同在模型内部并非行为上可互换:在具有相同错误答案的匹配项目上,因果干预可以有选择地抑制其中一种,而不对另一种产生同等影响。在三个开放权重家族中,移除拟合的来源方向使来源顺从性降低65-80个百分点,而移除用户或助手方向的效果则小得多,且移除用户方向显示出相反的偏好。一项基于来源与用户线索激活分别拟合的干预措施,在保持提示文本不变的情况下,使顺从性向两个方向移动。在琐事上拟合的权威方向也能无需重新拟合而迁移至PIQA和多轮SYCON对话,且移除该方向在五个家族中的四个里使错误来源顺从性降低数十个百分点,而在我们的评估规模下,MMLU-Pro或GSM8K准确率未检测到变化。因此,来源遵从与用户认同需要分开评估。

英文摘要

Language models tend to agree with whatever a user asserts, and post-training increasingly targets this sycophancy so that models evaluate claims on their merits rather than deferring to the user. Yet the same models are far more compliant when a wrong answer is attributed to a verified source, which is how retrieval results, tool outputs, and grounded-search content often present information. We measure this gap across five open-weight families and three closed APIs. A single verified-source note endorsing a wrong answer flips 45-88% of baseline-correct responses in seven of eight models, and compliance rises with how authoritative the note sounds. Source deference and user agreement are not behaviorally interchangeable inside the model: on matched items with the same wrong answer, causal interventions can selectively suppress one without equally affecting the other. In three open-weight families, removing a fitted source direction lowers source compliance by 65-80 percentage points while removing a user or assistant direction has far smaller effects, and removing the user direction shows the reverse preference. A separately fitted intervention derived from source-versus-user cue activations moves compliance in both directions while leaving the prompt text unchanged. An authority direction fitted on trivia also transfers to PIQA and multi-turn SYCON dialogues without refitting, and removing it lowers wrong-source compliance by tens of percentage points in four of five families with no detected change in MMLU-Pro or GSM8K accuracy at our evaluation sizes. Source deference and user agreement therefore need separate evaluation.

CommentsAccepted at NeurIPS 2026 (Main Conference, Poster). 33 pages, 8 figures. Project page: https://authority-bias.vercel.app/ . Code: https://github.com/Lossfunk/authority-bias

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑