AI 中文总结
针对通信受限分布式多智能体系统的任务调度问题,本文提出基于MDGAM的神经调度框架,结合GRMAPG算法,经实验验证其可提升任务完成性能。
AI 中文摘要
通信受限分布式多智能体系统中的协同任务调度极具挑战性,因为每个智能体必须基于部分且动态的观测结果做出决策,同时还要满足复杂的实际约束。现有启发式方法依赖手工设计的投标规则和重复共识,而许多基于学习的方法假设存在全局观测,且缺乏基于通信的显式协调。为解决这些局限,本文提出一种用于分布式多机器人任务分配(MRTA)的神经调度框架,包含多解码器图注意力模型(MDGAM)策略模型和无评论者的组相对多智能体策略梯度(GRMAPG)训练算法。MDGAM采用扩展图注意力机制联合更新节点和边特征,使用多个解码器生成任务选择决策与通信消息;GRMAPG从等价任务规划实例中构建组相对优势,替代传统多智能体强化学习(MARL)算法中使用的评论者网络,从而降低训练难度并提升收敛性能。在不同问题规模和通信范围下开展的实验表明,所提方法相比现有启发式方法和基于学习的方法提升了任务完成性能,消融实验、复杂度测试与泛化测试进一步验证了所提创新的有效性。
英文摘要
Cooperative task scheduling in communication-constrained distributed multi-agent systems is challenging because each agent must make decisions from partial and dynamic observations while satisfying complex practical constraints. Existing heuristics rely on handcrafted bidding rules and repeated consensus, whereas many learning-based methods assume global observations and lack explicit communication-based coordination. To address these limitations, this paper proposes a neural scheduling framework for distributed multi-robot task allocation (MRTA), consisting of a multi-decoder graph attention model (MDGAM) policy model and a critic-free group relative multi-agent policy gradient (GRMAPG) training algorithm. MDGAM uses an extended graph attention mechanism to jointly update node and edge features, and employs multiple decoders to generate task-selection decisions and communication messages. GRMAPG constructs group-relative advantages from equivalent task-planning instances to replace the critic network used in conventional MARL algorithms, thereby reducing training difficulty and improving convergence performance. Experiments under different problem scales and communication ranges show that the proposed method improves task-completion performance over existing heuristic and learning-based methods, while ablation, complexity, and generalization tests further validate the proposed innovations.