arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.12861cs.ROcs.AIcs.SYeess.SY

从简单奖励中揭示复杂的集体行为

Unveiling Complex Collective Behaviors from Simple Rewards

Yize Mi, Jianan Li, Liang Li, Shiyu Zhao

首次发表
浏览论文内容

中文总结 AI 辅助

研究多智能体强化学习中简单奖励与复杂集体行为的关系,提出两阶段EEC解释框架及智能体响应图,通过合作与竞争两个任务验证,揭示了MARL策略背后隐藏的几何结构。

中文摘要 AI 辅助

多智能体强化学习(MARL)在机器人集群中潜力巨大,但神经策略的黑箱性质使战略分析复杂化,限制了多机器人应用。此外,复杂的群体行为可从简单奖励中意外出现,揭示其背后机制至关重要。本文提出两阶段EEC(\LinkIII)解释框架,包括名为智能体响应图(ARM)的新分析工具,它揭示智能体在空间中的决策模式及聚集和回避区域。通过合作多机器人形状组装和竞争捕食者 - 猎物追逐逃避两个任务验证发现。在合作任务中,ARM确定未被占据的目标内部为机器人导航的期望目的地,随着中心被占据,目标区域自动向边界转移;在竞争任务中,ARM确定捕食者Voronoi图的边界为猎物智能体的收敛目的地。这两个任务证明了ARM发现机器人集群中MARL策略背后隐藏几何结构的能力。

英文摘要

Multi-agent Reinforcement Learning (MARL) holds great potential for robot swarms, but the black-box nature of neural policies complicates strategic analysis, limiting multi-robot applications. Furthermore, complex swarm behaviors can surprisingly emerge from simple rewards without explicit aggregation incentives. Unveiling the mechanisms behind this emergence is critical, but the disconnection between simple rewards and collective behaviors exacerbates interpretability challenges. This paper aims to reveal the hidden mechanisms in this process. We propose a two-stage EEC (\LinkIII) explanatory framework. This includes a novel analytical tool called the Agent Response Map (ARM), which reveals agents' decision-making patterns across space and identifies regions of aggregation and avoidance. ARM reveals that the robots implicitly learn the geometric fields of the environment and utilize these structures as desired targets for coordinated movement. We validate this finding across two distinct tasks: a cooperative multi-robot shape assembly and a competitive predator-prey pursuit-evasion. 1) In the cooperative task, ARM identifies the unoccupied target interior as the desired destination for robot navigation. As the center becomes occupied, this target region automatically shifts toward the boundary, demonstrating the robots' capacity to autonomously explore unoccupied areas. 2) In the competitive task, ARM surprisingly identifies the boundary of the predators' Voronoi diagram as the convergence destination for prey agents. Together, these two tasks demonstrate the capability of ARM to discover the hidden geometric structures underlying MARL policies in robot swarms.

发表机构

  • College of Computer Science and Technology, Zhejiang University(浙江大学计算机科学与技术学院)
  • WINDY Lab, School of Engineering, Westlake University(西湖大学工程学院风语实验室)
  • Shanghai AI Laboratory(上海人工智能实验室)
  • Max Planck Institute of Animal Behavior(马克斯·普朗克动物行为研究所)
  • Department of Computer and Information Science, University of Konstanz(康斯坦茨大学计算机与信息科学系)
  • Centre for the Advanced Study of Collective Behaviour, University of Konstanz(康斯坦茨大学集体行为高级研究中心)
  • Department of Biology, University of Konstanz(康斯坦茨大学生物系)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑