arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

合作的涌现:一种声誉调节的强化学习

Emergence of cooperation: A reputation-modulated reinforcement learning

Chenyang Zhao, Jiqiang Zhang, Li Chen, Yong Zou

arXiv 2608.20016首次发表:更新:

AI 中文总结

本研究提出基于强化学习的空间囚徒困境博弈,让配备Q学习的智能体通过局部声誉指标整合信息,发现声誉调节学习可显著促进合作涌现,还揭示了合作与背叛间的不连续相变机制。

AI 中文摘要

声誉被广泛认为是维持合作的关键机制,但大多数现有的博弈论模型主要将声誉视为调节收益、交互结构或策略更新规则的外部因素。然而在许多社会情境中,声誉主要作为一种信息发挥作用——它塑造个体对自身经历的解读方式及对他人行为的评估方式。为填补这一空白,我们提出一种基于强化学习范式的空间囚徒困境博弈,其中配备Q学习的智能体通过局部定义的声誉指标整合个体与社会信息以指导决策。研究结果显示,声誉调节学习显著促进合作行为的涌现,且随着诱惑值增大,我们观测到从完全合作到完全背叛的不连续相变。合作通过合作集群的成核过程扩散,而这些集群的解体则驱动系统进入完全背叛的吸收态。总体而言,本研究表明声誉促进合作不仅通过提供直接激励,还通过重塑智能体学习与适应所依赖的社会信息环境实现。

英文摘要

Reputation is widely recognized as a key mechanism for sustaining cooperation. However, most existing game-theoretic models treat reputation primarily as an external factor that modulates payoffs, interaction structures, or strategy update rules. In many social contexts, though, reputation operates primarily as information -- it shapes how individuals interpret their own experiences and assess the behavior of others. To bridge this gap, we propose a spatial prisoner's dilemma game grounded in the reinforcement learning paradigm, in which agents equipped with Q-learning integrate both individual and social information via a locally defined reputation metric to guide their decisions. Our results reveal that reputation-modulated learning significantly promotes the emergence of cooperative behavior, and we observe a discontinuous phase transition from full cooperation to full defection as the temptation increases. Cooperation spreads through the nucleation of cooperative clusters, whereas the disintegration of these clusters drives the system into an absorbing state of complete defection. Overall, this study demonstrates that reputation facilitates cooperation not only by providing direct incentives but also by reshaping the social information landscape that agents rely on for learning and adaptation.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑