arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.07662cs.CLcs.AIcs.HC

AI模型如何管理认知权威:对用户分歧回应的分类与比较分析

How AI Models Manage Epistemic Authority: A Taxonomy and Comparative Analysis of Responses to User Disagreement

  • University of Edinburgh(爱丁堡大学)
  • Middle East Technical University(中东技术大学)

机构由 AI 辅助整理,请以论文原文为准。

Riyadh Alnasser, Yusuf Mücahit Çetinkaya, Sumin Zhao, Tuğrulcan Elmas

中文总结 AI 辅助

本研究提出六类挑战类型和四层分析框架,构建数据集分析14个模型对用户分歧的回应,发现模型在验证用户与维持主张间存在矛盾,且权威转移在建议任务中更常见。

中文摘要 AI 辅助

大型语言模型越来越多地被用作建议和信息来源,包括在高风险场景中,但人们对它们如何回应用户分歧知之甚少。我们研究模型在用户质疑其答案时如何管理其认知权威,这里指的是其对知识、能力或建议权的声称。基于会话分析,我们引入了一个包含六种挑战类型的分类法,以及一个用于分析每种回应的四层框架:原始主张是否被维持或改变、权威位于何处、分歧如何被社交性地管理,以及提供了何种证据支持。我们构建了一个新的数据集,包含来自14个模型的2,310个受控挑战场景和32,340个对应回应,并使用我们的框架通过LLM-as-judge流水线进行分析,为未来的评估和基准设计提供了词汇表。我们发现模型表现出矛盾的行为:在85%的回应中验证用户,但在65%的回应中维持其原始主张。在33%的回应中明确道歉,但其中59%的道歉伴随着对原始主张的维持。它们在建议任务中最常转移权威,在28%的回应中如此,在健康建议中达到57%,在法律建议中达到49%,而在事实任务中为6%,在解释任务中为3%。放弃原始主张的比例从GPT-5.2的0.8%到DeepSeek 7B的40%不等,而完全替换原始主张的情况总体罕见,为1.5%。

英文摘要

Large language models are increasingly used as sources of advice and information, including in high-stakes settings, yet little is known about how they respond to user disagreement. We study how a model manages its epistemic authority, referring here to its claim to knowledge, competence, or the right to advise, once a user challenges its answer. Building on Conversation Analysis, we introduce a taxonomy of six challenge types and a four-layer framework for analysing each response: whether the original claim is maintained or changed, where authority is located, how the disagreement is socially managed, and what kind of evidential support is offered. We construct a new dataset of 2,310 controlled challenge scenarios and 32,340 corresponding responses from 14 models, and analyse them using our framework with an LLM-as-judge pipeline, providing a vocabulary which future evaluation and benchmark design can build on. We find that models show conflicting behaviour: they validate users in 85% of responses but maintain their original claim in 65%. They explicitly apologise in 33% of responses, yet 59% of those apologies accompany maintenance of the original claim. They transfer authority most often in advice tasks, doing so in 28% of responses and reaching 57% in health advice and 49% in legal advice, compared with 6% in fact and 3% in explanation tasks. Abandonment of the original claim ranges from 0.8% for GPT-5.2 to 40% for DeepSeek 7B, while complete replacement of the original claim is rare overall at 1.5%.

补充信息

↑