arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

基于算法利他主义的异构多机器人系统中的协作风险感知探索

Cooperative Risk-Aware Exploration in Heterogeneous Multi-Robot Systems Using Algorithmic Altruism

Brooks A. Butler, Jair Certório, João P. Hespanha, Magnus Egerstedt

arXiv 2608.28409首次发表:更新:

发表机构

School of Electrical and Computer Engineering at Oklahoma State University; Instituto Tecnológico de Aeronáutica; University of California, Santa Barbara; University of North Carolina, Chapel Hill(俄克拉荷马州立大学电气与计算机工程学院; 航空技术研究所; 加州大学圣巴巴拉分校; 北卡罗来纳大学教堂山分校)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究针对异构多机器人系统,提出基于算法利他主义的博弈论框架,通过社会纳什均衡调整效用实现风险合理分配,减少冗余探索并提升团队性能,经仿真与硬件实验验证有效。

AI 中文摘要

多机器人系统非常适合在危险环境中执行探索任务,但要实现有效部署,不仅需要确定机器人应在何处收集信息,还需要在异构团队成员间合理分配风险。本文基于受生态学启发的利他行为,开发了一种用于协作风险感知探索的博弈论框架。每个机器人选择有限时间范围的轨迹,以最大化信息增益,同时惩罚冗余探索和预期风险暴露。通过智能体特定的价值参数引入异构性,这些参数用于编码利他耦合,该耦合通过受汉密尔顿法则启发的亲缘权重进行建模。我们引入了一种轨迹规划的博弈论结构,定义了社会纳什均衡,该均衡会根据智能体的亲缘关系调整智能体行动的效用。这种效用塑造使智能体将自身轨迹选择对队友的影响内化,鼓励低价值机器人在这样做有利于高价值智能体并提升团队性能时接受风险。我们为智能体定义了一种探索效用,其奖励区域覆盖和不确定性减少,同时惩罚冗余和风险,从而能够在后退时域规划器中进行基于投影梯度的航点优化。仿真结果表明,利他规划减少了冗余探索,改善了机器人间的间距,并根据智能体价值重新分配风险,同时保持了相当的地图覆盖范围。我们还在硬件实验中验证了该方法,轮式机器人使用单积分控制器和障碍证书跟踪规划的航点。

英文摘要

Multi-robot systems are well-positioned for exploration in hazardous environments, but effective deployment requires deciding not only where robots should gather information, but also how risk should be distributed across heterogeneous team members. This paper develops a game-theoretic framework for cooperative risk-aware exploration based on ecologically inspired altruistic behavior. Each robot selects a finite-horizon trajectory to maximize information gain while penalizing redundant exploration and expected hazard exposure. Heterogeneity is introduced through agent-specific value parameters for encoding altruistic coupling, which is modeled through relatedness weights inspired by Hamilton's rule. We introduce a game-theoretic structure for trajectory planning that defines a Social Nash Equilibrium, which modifies the utility of agent actions according to agent relatedness. This utility shaping causes agents to internalize the effect of their trajectory choices on teammates, encouraging lower-valued robots to accept risk when doing so benefits higher-valued agents and improves team performance. We define an exploration utility for agents that rewards area coverage and uncertainty reduction, while also penalizing redundancy and risk, enabling projected gradient-based waypoint optimization in a receding-horizon planner. Simulations show that altruistic planning reduces redundant exploration, improves inter-robot separation, and reallocates risk according to agent value while maintaining comparable map coverage. We further demonstrate the approach in hardware experiments, where planned waypoints are tracked by wheeled robots using single-integrator controllers and barrier certificates.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑