arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.19555cs.MA

CHMAS:一种用于多智能体强化学习的耦合分层框架

CHMAS: A Coupled Hierarchical Framework for Multi-Agent Reinforcement Learning

Dongming Wang, Jie Xu, Yanyu Zhang, Wei Ren

首次发表
浏览论文内容

中文总结 AI 辅助

本文针对多智能体强化学习跨时间尺度平衡协调与执行的挑战,提出CHMAS框架,将决策分解为集中战略规划与分布式战术执行,有双向信息流和反馈机制,还给出异步更新协议,经理论分析与实验验证,该框架有效。

中文摘要 AI 辅助

多智能体强化学习(MARL)系统在跨不同时间尺度平衡全局协调与局部执行方面面临根本挑战。本文介绍了耦合分层多智能体系统(CHMAS),这是一个新颖的框架,它将多智能体决策分解为具有双向信息流的集中式战略规划和分布式战术执行。战略层整合所有智能体状态与唯一全局环境状态,每\(T\)时间步生成指导行动,战术智能体执行由战略指导和局部邻域观测增强的分布式策略。与现有单向控制的分层方法不同,CHMAS建立了反馈机制,累积战术奖励通过耦合系数\(\lambda\)影响战略目标。为解决分层学习中的非平稳性,提出异步更新协议,战略参数每\(N_f\)个战术情节更新一次。给出了捕获完整系统动态的一般双层公式和便于严格分析的可处理加法近似。理论分析证明在标准假设下,经过\(K\)次战略更新后,该异步方案使战略层达到\(\mathcal{O}(\log K/\sqrt{K})\)收敛。在多智能体觅食领域的实验验证表明成功学习了空间分区探索策略,尽管存在分层耦合,两层仍稳定收敛。

英文摘要

Multi-agent reinforcement learning (MARL) systems face fundamental challenges in balancing global coordination with local execution across different temporal scales. This paper introduces the Coupled Hierarchical Multi-Agent System (CHMAS), a novel framework that decomposes multi-agent decision-making into centralized strategic planning and distributed tactical execution with bidirectional information flow. The strategic layer integrates all agents' states with an exclusive global environmental state to generate guidance actions every $T$ timesteps, while tactical agents execute distributed policies augmented by strategic guidance and local neighborhood observations. Unlike existing hierarchical approaches with unidirectional control, CHMAS establishes a feedback mechanism where accumulated tactical rewards influence strategic objectives through a coupling coefficient $λ$, ensuring strategic plans remain grounded in tactical feasibility. To address the non-stationarity inherent in hierarchical learning, we propose an asynchronous update protocol where strategic parameters update every $N_f$ tactical episodes, allowing tactical policies to converge to quasi-stationary points between strategic changes. We present both a general bi-level formulation capturing full system dynamics and a tractable additive approximation enabling rigorous analysis. Theoretical analysis proves that this asynchronous scheme achieves $\mathcal{O}(\log K/\sqrt{K})$ convergence for the strategic layer after $K$ strategic updates under standard assumptions. Experimental validation in a multi-agent foraging domain demonstrates successful learning of spatially partitioned exploration strategies, with both layers converging stably despite hierarchical coupling.

补充信息

↑