arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.03958cs.AI

基础模型的博弈论:通过相似性推理实现理性合作的新路径

A game theory for foundation models shows new paths to rational cooperation through similarity inference

Alexander Meulemans, Maciej Wołczyk, Marissa A. Weis, Rajai Nasser, Roberta Rocca, Seijin Kobayashi, Guillaume Lajoie, Angelika Steger, Blake Richards, Marcus H… 展开作者

Alexander Meulemans, Maciej Wołczyk, Marissa A. Weis, Rajai Nasser, Roberta Rocca, Seijin Kobayashi, Guillaume Lajoie, Angelika Steger, Blake Richards, Marcus Hutter, James Manyika, Rif A. Saurous, João Sacramento, Blaise Agüera y Arcas

首次发表
浏览论文内容

中文总结 AI 辅助

本文提出针对基础模型智能体的嵌入贝叶斯智能体理论模型,通过嵌入均衡机制实现相似性推理,破解经典博弈论预测的社会困境中相互背叛问题,达成稳定合作。

中文摘要 AI 辅助

随着由基础模型驱动的自主智能体日益融入社会与经济系统,理解支配其集体行为的原则对确保安全与合作至关重要。经典博弈论是建模理性互动的主流框架,其建立在“解耦智能体”假设之上,即智能体将自身决策视为独立于环境与其他行为者的存在。然而,现代AI智能体在预测自身未来行动的同时,也会对外部观测结果进行预测。本文报告了一项引人注目的发现:在风格化社会困境中互动时,进行最优规划的基础模型智能体始终会收敛至稳定合作状态,这与经典博弈论中相互背叛的预测直接矛盾。为理解该现象,本文提出“嵌入贝叶斯智能体”这一针对基础模型智能体的理论模型。通过从解耦智能体转向嵌入智能体,这些智能体将自身建模为所处宇宙的一部分,并对自身决策算法保持认知不确定性。研究表明,嵌入智能体可通过推断他人行为是否相似,将自身规划过程中的思考作为证据:合作决策可预测相似伙伴的类似决策。本文通过“嵌入均衡”形式化了这种相似性推理机制,该机制是替代纳什均衡的新型解概念,为现代AI智能体的社会行为提供了基础博弈论。

英文摘要

As autonomous agents powered by foundation models are increasingly integrated into social and economic systems, understanding the principles governing their collective behavior is essential for ensuring safety and cooperation. Classical game theory, the dominant framework for modeling rational interaction, is built upon the assumption of `decoupled agency,' where agents treat their own decision-making as independent of the environment and other actors. Modern AI agents, however, jointly predict their own future actions alongside external observations. Here, we report a striking finding: when interacting in stylized social dilemmas, foundation model agents engaging in optimal planning consistently converge to stable cooperation, directly contradicting classical game-theoretic predictions of mutual defection. To understand this phenomenon, we introduce the `embedded Bayesian agent,' a theoretical model for foundation model agents. By shifting from decoupled to embedded agency, these agents model themselves as part of the universe they inhabit, maintaining epistemic uncertainty about their own decision-making algorithms. We show that by inferring whether others are behaviorally similar, an embedded agent treats its own deliberation during planning as evidence: a decision to cooperate predicts a similar decision by a similar partner. We formalize this mechanism of similarity inference through the `embedded equilibrium,' a novel solution concept replacing the Nash equilibrium to provide a foundational game theory for the social behavior of modern AI agents.

补充信息

↑