arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2412.12326cs.MAcs.AIcs.LG

通过建议共享在多智能体强化学习中实现集体福利

Achieving Collective Welfare in Multi-Agent Reinforcement Learning via Suggestion Sharing

  • Warwick Manufacturing Group, University of Warwick, Coventry, UK(沃里克制造集团,沃里克大学,科文特里,英国)
  • School of Electrical Engineering and Computer Science, Louisiana State University, USA(电气工程与计算机科学学院,路易斯安那州立大学,美国)
  • Department of Statistics, University of Warwick, Coventry, UK(统计学系,沃里克大学,科文特里,英国)
  • Alan Turing Institute, London, UK(艾伦·图灵研究所,伦敦,英国)

机构由 AI 辅助整理,请以论文原文为准。

Yue Jin, Shuangqing Wei, Giovanni Montana

更新

AI总结:

针对多智能体强化学习中个体利益与集体利益冲突的问题,提出一种通过交换行动建议实现协作的方法,泄露隐私更少且无需设计内在奖励,性能与现有基线相当并得到理论支撑。

AI中文摘要:

在人类社会中,自利与集体福祉之间的冲突往往会阻碍实现共同福利的努力。公地悲剧、社会困境等相关概念在我们的日常生活中屡见不鲜。随着人工智能体越来越多地充当人类的自主代理,我们提出了一种新颖的多智能体强化学习(MARL)方法来解决这一问题——即使个体智能体的利益与集体利益相冲突,也能学习最大化集体回报的策略。传统的协作式多智能体强化学习解决方案涉及共享奖励、价值和策略,或设计内在奖励来鼓励智能体学习集体最优策略,与之不同,我们提出了一种新颖的多智能体强化学习方法,让智能体之间交换行动建议。与共享奖励、价值或策略相比,我们的方法泄露的隐私信息更少,同时无需设计内在奖励就能实现有效协作。我们的理论分析为该算法提供了支撑,分析确立了集体目标与个体目标之间差异的界,阐明了共享建议如何能使智能体的行为与集体目标保持一致。实验结果表明,我们的算法性能与依赖价值共享、策略共享或内在奖励的基线方法相当。

英文摘要:

In human society, the conflict between self-interest and collective well-being often obstructs efforts to achieve shared welfare. Related concepts like the Tragedy of the Commons and Social Dilemmas frequently manifest in our daily lives. As artificial agents increasingly serve as autonomous proxies for humans, we propose a novel multi-agent reinforcement learning (MARL) method to address this issue - learning policies to maximise collective returns even when individual agents' interests conflict with the collective one. Unlike traditional cooperative MARL solutions that involve sharing rewards, values, and policies or designing intrinsic rewards to encourage agents to learn collectively optimal policies, we propose a novel MARL approach where agents exchange action suggestions. Our method reveals less private information compared to sharing rewards, values, or policies, while enabling effective cooperation without the need to design intrinsic rewards. Our algorithm is supported by our theoretical analysis that establishes a bound on the discrepancy between collective and individual objectives, demonstrating how sharing suggestions can align agents' behaviours with the collective objective. Experimental results demonstrate that our algorithm performs competitively with baselines that rely on value or policy sharing or intrinsic rewards.

补充信息

↑