arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

大型语言模型表现出类似人类的贝叶斯虚伪

Large Language Models Exhibit Human-Like Bayesian Hypocrisy

Nykko Vitali, Mahzarin R. Banaji

arXiv 2609.35779首次发表:更新:

发表机构

Harvard University(哈佛大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究发现LLM在贝叶斯推理上接近人类水平,但同样表现出谴责他人使用相同规则的虚伪,凸显自我与他人判断的分离。

AI 中文摘要

鉴于大型语言模型(LLM)近期取得的成就,前沿模型预计在贝叶斯推理任务上表现良好,至少与人类相当。此外,没有理由预期LLM会谴责那些提供同样贝叶斯判断的他人,而这种缺陷在人类决策中已被观察到(Cao等人,2019)。在包含48个实验条件、超过5000次试验的5项实验中,GPT-4o和Claude 3.7 Sonnet在贝叶斯推理任务的两种变体上接受了测试。我们还评估了LLM对提供与自身相同推理任务的一个假设人物的能力和道德的评价。LLM在贝叶斯任务上的表现接近人类水平,尽管其推理更加基于规则且僵化。令人惊讶的是,与人类相似但程度更高,LLM也表现出同样的虚伪,即谴责那些像它们一样运用了贝叶斯规则的他人。通过展示贝叶斯虚伪,LLM凸显了一种类似人类的错误,即自我表现与他人判断之间的分离,并警示在统计保真度与公平规范相冲突的领域中应谨慎使用它们。

英文摘要

Given recent achievements of large language models (LLMs), frontier models are expected to perform well on Bayesian reasoning tasks, at least as well as humans. Furthermore, there is no reason to expect that LLMs will condemn others who offer those very same Bayesian judgments, a fallibility observed in human decision-making (Cao, et al., 2019). In 5 experiments with 48 experimental conditions employing over 5,000 trials, GPT-4o and Claude 3.7 Sonnet were tested on two variations of a Bayesian reasoning task. We also assessed LLM evaluation of the competence and morality of a hypothetical person who had offered the same reasoning task as them. LLMs hovered near human performance on the Bayesian task, though their reasoning was more rule-based and rigid. Surprisingly, like humans but to a greater extent, LLMs also demonstrated the same hypocrisy in condemning others who, like them, had deployed Bayes' rule. In demonstrating Bayesian hypocrisy, LLMs highlight a humanlike error of a dissociation between self-performance and other-judgment, and caution against their use in domains where statistical fidelity and fairness norms collide.

CommentsMain text (38 pages) with supplementary materials appended (213 pages total). Preregistered with data and analysis code at https://osf.io/95tfc

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑