arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.09258cs.RO

面向任务的编队决策:通过强化学习实现对攻击型集群的驱赶

Task-Oriented Formation Decision via Reinforcement Learning: Herding an Attacking Swarm

Zhaozong Wang, Guibin Sun, Jinyong Chen, Rui Zhou

首次发表
浏览论文内容

中文总结 AI 辅助

本文针对攻击型集群驱赶任务,提出基于强化学习的编队决策方法,通过参数化编队形状缓解防御者机动性劣势,在仿真与物理平台上验证了方法的可行性与扩展性。

中文摘要 AI 辅助

多机器人系统可通过组织成特定任务的编队完成单机器人难以胜任的任务。与现有多机器人形状编队研究不同,本文聚焦于驱赶任务,研究面向任务的编队决策问题。该任务极具挑战性,因为攻击者(attackers)具备更优的机动性且策略未知。为应对这些挑战,本文提出两项创新成果:其一,采用低维参数向量对编队形状进行编码,这种参数化表示将编队决策重构为参数优化问题,解决了预定义形状灵活性有限的问题;通过优化编队参数,防御者(defenders)可生成持续适配任务需求的编队形状,从而缓解机动性劣势。其二,开发基于强化学习的策略来调控编队参数,该策略在涵盖多种攻击策略的仿真环境中进行离线训练,学习到的策略可在在线部署时有效应对对抗性不可预测性。与三种基准方法的对比仿真表明,本文方法可成功完成极具挑战性的驱赶任务;额外的可扩展性仿真进一步验证了其适用于包含数十个机器人的仿真场景;本文还在由3个攻击者和7个防御者组成的物理机器人平台上验证了所提方法的实际可行性。

英文摘要

Multi-robot systems can accomplish tasks that are difficult for a single robot by organizing into task-specific formations. Different from existing studies on multi-robot shape formation, we here study the task-oriented formation decision problem, with a focus on the herding task. This task is challenging due to the attackers' superior maneuverability and their unknown strategies. To address these challenges, we propose the following novel results. First, we encode the formation shape using a low-dimensional parameter vector. This parametric representation reformulates the formation decision as a parameter optimization problem, thereby resolving the limited flexibility of predefined shapes. By optimizing these formation parameters, the defenders' maneuverability disadvantage is mitigated through a formation shape that continuously adapts to task requirements. Second, we develop a reinforcement learning-based policy to regulate the formation parameters. Trained offline in simulations covering diverse attacking strategies, the learned policy can effectively handle adversarial unpredictability during online deployment. Comparative simulations against three baselines demonstrate that our method can successfully accomplish challenging herding tasks. Additional scalability simulations further verify its applicability to simulated scenarios involving dozens of robots. We also validate the practical feasibility of our method on a physical robotic platform with 3 attackers and 7 defenders.

补充信息

↑