arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

网络拓扑与对手信息在多智能体强化学习系统中塑造合作的作用

The Role of Network Topology and Opponent Information in Shaping Cooperation in Multi-Agent Reinforcement Learning Systems

Seongho Son, Stephen Hailes, Mirco Musolesi

arXiv 2608.28977首次发表:更新:

发表机构

University College London; UCL Centre for Artificial Intelligence; University of Bologna(伦敦大学学院; 伦敦大学学院人工智能中心; 博洛尼亚大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究探究多智能体强化学习中,图拓扑、对手行动历史与身份信息对迭代囚徒困境博弈合作行为的影响,发现节点邻居数、平均路径长度及伙伴选择等因素对合作的作用。

AI 中文摘要

已有多项研究探讨了图拓扑对人工智能体间合作的影响,但多数文献聚焦于通过策略模仿建模智能体的适应过程,这类模仿仅依赖于其他智能体的累积收益。本文研究的场景中,每个智能体使用深度强化学习学习进行双人迭代囚徒困境(IPD)博弈,每个智能体被表示为图中的一个节点,其邻居构成了可与之交互的对手池。在每一轮IPD博弈中,智能体被提供关于对手的不同类型信息,包括行动历史和对手身份。对不同图拓扑的实验结果显示,每个节点的邻居数量和平均路径长度是影响合作出现的主要因素。我们还表明,尽管伙伴选择通过限制对手池的多样性促进了相互合作,但向智能体提供对手身份会阻碍合作策略的扩散。

英文摘要

Several works have investigated the influence of graph topology on cooperation among artificial agents, while the majority of the literature has focused on modelling agents' adaptation through strategy imitation, which relies solely on the cumulative payoffs of others. This paper investigates scenarios in which each agent learns to play the two-player Iterated Prisoner's Dilemma (IPD) using deep reinforcement learning. Each agent is represented as a node in a graph, where its neighbours constitute the pool of opponents with whom it can interact. During each IPD episode, agents are provided with different types of information about their opponent, consisting of action history and opponent identity. Experimental results across different graph topologies show that the number of neighbours per node and the average path length are the main factors affecting the emergence of cooperation. We also show that, while partner selection fosters mutual cooperation by limiting the diversity of the opponent pool, providing agents with the identity of their opponent hinders the proliferation of cooperative strategies.

Comments19 pages, 13 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑