arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

在变化的世界中学习合作:对未来关注如何促进跨尺度合作

Learning to cooperate in a changing world: How caring about the future promotes cooperation across scales

Yuxin Geng, Xingru Chen, Xin Wang, Hongwei Zheng, Longzhao Liu, Shaoting Tang, Feng Fu

arXiv 2609.34005首次发表:更新:

发表机构

Beihang University; Zhongguancun Laboratory; Hangzhou International Innovation Institute, Beihang University; Key Laboratory of Mathematics, Informatics and Behavioral Semantics, Beihang University; Beijing Advanced Innovation Center for Future Blockchain and Privacy Computing, Beihang University; Beijing Academy of Blockchain and Edge Computing(北京航空航天大学; 中关村实验室; 北京航空航天大学杭州创新研究院; 北京航空航天大学数学、信息学与行为语义重点实验室; 北京航空航天大学未来区块链与隐私计算高级创新中心; 北京区块链与边缘计算研究院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究通过多智能体强化学习,推导出跨尺度稳定合作的解析条件,发现重视未来能引导自利个体合作,为合作型AI提供统一框架。

AI 中文摘要

在社会困境中,个体需要放弃短期诱惑,通过合作实现协同性的集体成果。以往的研究考察了合作得以演化的机制,包括直接互惠、间接互惠、环境随机性、网络互惠和人口随机性。这些机制大多在自然选择或社会学习背景下被研究。作为同一枚硬币的另一面,研究自我学习(即个体基于自身经验进行适应)下的合作同样重要。在此,我们聚焦于多智能体强化学习的动力学,并推导出这些机制在跨尺度的学习动力学下能够稳定合作的解析条件。我们发现,当智能体重视未来而非短期诱惑时,强化学习能够引导自利个体走向合作。我们的工作提供了一种统一的方法来识别在变化世界中学习合作的决定因素,从而为合作型人工智能的发展铺平道路。

英文摘要

In social dilemmas, individuals need to forgo short-term temptations to achieve synergistic collective outcomes through cooperation. Previous work has examined mechanisms through which cooperation can evolve, including direct reciprocity, indirect reciprocity, environmental stochasticity, network reciprocity, and demographic stochasticity. These mechanisms have largely been studied under natural selection or social learning, where strategies with higher payoffs are more likely to spread. It is equally important to study cooperation under self-learning, where individuals adapt through their own experiences. Here, we focus on multi-agent reinforcement learning and derive analytical conditions under which these mechanisms stabilize cooperation under learning dynamics across scales. We find that reinforcement learning can steer self-interested individuals toward cooperation when they value the future over short-term temptation. Our work provides a unified approach to identifying the determinants of learning to cooperate in a changing world, thereby paving the way for the advancement of cooperative artificial intelligence.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑