arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI智能体中的贝叶斯推理与动机推理

Bayesian and Motivated Reasoning in AI Agents

Eddie Yang

arXiv 2608.00339首次发表:更新:

AI 中文总结

该研究发现AI智能体的结论受先验信念影响,相同数值数据因框架不同会得出不同结论,揭示了委托其决策存在依赖未指定先验的风险。

AI 中文摘要

AI智能体越来越多地在其结论可指导重要决策的场景中执行开放式任务。我们提供证据表明,当实质性框架发生变化时,AI智能体从相同的数值数据中会得出不同的结论。我们通过固定证据,仅改变证据出现的场景,在医学、选举取证和地缘政治预测等高风险领域展示了这种行为。在12个智能体-领域对比中,智能体的结论受其先验信念的强烈影响:当结论围绕智能体已认为可能的命题构建时,它们更可能得出肯定结论;而当框架与它们的先验信念冲突时,则相反。框架还改变了部分智能体的工作方式:它们会进行更广泛的搜索、选择不同的分析规格,并对相同证据的评估方式不同。这些结果指出了将决策委托给AI智能体的特定风险,因为它们的决策可能依赖于任务中未指定且决策记录中不可见的先验信念。

英文摘要

AI agents increasingly perform open-ended tasks in settings where their conclusions can guide consequential decisions. We provide evidence that AI agents draw different conclusions from identical numerical data when the substantive framing changes. We demonstrate this behavior in high-stakes domains in medicine, election forensics, and geopolitical forecasting by holding the evidence fixed while changing the scenario in which the evidence appears. Across twelve agent-domain comparisons, agents' conclusions are strongly influenced by their prior beliefs. They are more likely to reach an affirmative conclusion when it is framed around a proposition they already regard as likely, while the reverse holds when the framing conflicts with their prior. The framing also changes how some agents work: they search more extensively, choose different analytical specifications, and evaluate the same evidence differently. These results identify a particular risk of delegating decision-making to AI agents, as their decisions may depend on prior beliefs that are neither specified in the task nor visible in the decision record.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑