arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

进化稳定性并不保证学习可达性:多智能体强化学习视角下的合作涌现

Evolutionary Stability Does Not Guarantee Learning Accessibility: A Multi-Agent Reinforcement Learning Perspective on Cooperation Emergence

Yijie Wang

arXiv 2609.27664首次发表:更新:

发表机构

Liupanshui Normal University(六盘水师范学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究通过三智能体治理博弈证明,进化稳定性不保证多智能体强化学习中的合作可达性,并提出了比较群体稳定性与有限样本学习可达性的框架。

AI 中文摘要

合作涌现是多智能体系统中的核心问题,因为去中心化的智能体必须在适应其他智能体行为变化的同时进行协调。进化博弈论识别出策略上稳定的结果,但在群体调整动态下的稳定性并不一定意味着有限样本的学习智能体能够通过局部奖励反馈达到相同的结果。我们在一个透明的、由治理动机驱动的三智能体博弈中研究这一区别,该博弈涉及政府、平台企业和用户。我们推导了固定阶段博弈激励下的复制者动态,在对称初始条件网格上评估合作进化盆地,并将其与三种去中心化基于价值的学习器的学习盆地估计进行比较。学习分析使用独立Q学习(ε-贪心动作选择)、缩放玻尔兹曼探索和SA--EA BQL,在相同的支付环境和结果标准下进行。进化盆地在采样网格上的体积为$V_E=1.00$。经验学习盆地对于ε-IQL为$0.88$,对于缩放玻尔兹曼和SA--EA BQL均为$0.00$。诊断轨迹显示,更广泛的动作多样性和非零价值分离可以与在该固定配置中未能维持合作联合动作共存。这些结果表明,进化稳定性和学习可达性是耦合博弈-学习系统的不同属性。共享单车设置是一个激励应用;更广泛的贡献是一个框架,用于在指定的多智能体学习动态下比较群体水平稳定性与合作的有限样本可达性。

英文摘要

Cooperation emergence is a central problem in multi-agent systems because decentralized agents must coordinate while adapting to the changing behavior of others. Evolutionary game theory identifies strategically stable outcomes, but stability under a population adjustment dynamic need not imply that finite-sample learning agents can reach the same outcome through local reward feedback. We study this distinction in a transparent three-agent governance-motivated game involving a government, a platform firm, and users. We derive replicator dynamics for the fixed stage-game incentives, evaluate the cooperative evolutionary basin on a symmetric initial-condition grid, and compare it with learning-basin estimates for three decentralized value-based learners. The learning analysis uses independent Q-learning with $\varepsilon$-greedy action selection, scaled Boltzmann exploration, and SA--EA BQL under the same payoff environment and outcome criterion. The evolutionary basin has volume $V_E=1.00$ on the sampled grid. The empirical learning basin is $0.88$ for $\varepsilon$-IQL and $0.00$ for both scaled Boltzmann and SA--EA BQL. Diagnostic traces show that broader action diversity and nonzero value separation can coexist with failure to sustain the cooperative joint action in this fixed configuration. These results indicate that evolutionary stability and learning accessibility are distinct properties of a coupled game--learning system. The shared-bike setting is a motivating application; the broader contribution is a framework for comparing population-level stability with the finite-sample accessibility of cooperation under specified multi-agent learning dynamics.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑