CoRe-MARL:未知动态下基于循环多智能体强化学习的合作再分配
CoRe-MARL: Cooperative Redistribution Under Unknown Dynamics Using Recurrent Multi-Agent Reinforcement Learning
浏览论文内容
中文总结 AI 辅助
针对应急物资再分配中未知动态与独立决策问题,提出基于Dec-POMDP和循环MAPPO的合作多智能体框架,实现集中训练去中心化执行,有效缩小服务差距并提升最差区域服务。
中文摘要 AI 辅助
应急管理援助计划,如救济物资分配,对于向受灾社区提供必要物资至关重要。然而,这些计划在一个由地方中心组成的去中心化网络中运作,这些中心面临不确定的本地需求和供应动态,导致本地服务可用性不一致。在这些中心之间重新分配物资可以减少这些不平衡,但中心往往在信息有限且运输中断的情况下独立做出决策。本研究开发了CoRe-MARL,一个合作式多智能体强化学习(MARL)框架,通过制定去中心化部分可观测马尔可夫决策过程(Dec-POMDP)来实现。我们将每个中心视为一个智能体,学习再分配策略,以改善最差情况区域的服务并减少跨区域的服务差距,同时保护网络范围的服务。我们引入了一个循环网络,无需直接观测即可捕捉不断变化的供需动态,而多智能体近端策略优化(MAPPO)则实现了集中训练和去中心化执行(CTDE)。我们在一个具有多样化轨迹的模拟环境中评估该框架,其中演员和MAPPO评论家无法观测到精确动态。我们将循环MAPPO与循环独立PPO(IPPO)以及仅本地的启发式方法进行比较,发现MAPPO减少了地方中心之间的服务差距,并增强了服务最差中心的性能,同时保持了具有竞争力的网络范围服务。循环MAPPO在多样化轨迹模式中也表现出一致的性能,展示了其适应不断变化动态的能力。研究结果表明,合作学习能够实现去中心化再分配,并在不确定和不断变化的动态下提高服务的公平性。
英文摘要
Emergency management assistance programs, such as relief distribution, are essential for delivering necessary supplies to affected communities. However, these programs operate in a decentralized network of local centers that face uncertain local demand and supply dynamics, resulting in inconsistent avail- ability of local services. Redistribution of supplies among these local centers reduces these imbalances, but the centers often make decisions independently, with limited information and disrupted transportation. This study develops CoRe-MARL, a cooperative multi-agent reinforcement learning (MARL) framework, by formulating a decentralized partially observable Markov decision process (Dec-POMDP). We treat each center as an agent that learns a redistribution policy to improve the service in the worst-case region and reduce the service gap across regions while protecting network-wide service. We incorporate a recurrent network that captures evolving supply and demand dynamics without direct observation, while multi-agent proximal policy optimization (MAPPO) enables centralized training and decentralized execution (CTDE). We evaluate the framework in a simulated environment with diverse trajectories, where exact dynamics are not observed by actors and the MAPPO critic. We compare the recurrent MAPPO with the recurrent independent PPO (IPPO) and a local only heuristic, and find that MAPPO reduces the service gap across local centers and enhances service for the worst-served center while maintaining competitive network-wide service. The recurrent MAPPO also shows consistent performance across diverse trajectory patterns, demonstrating its ability to adapt to evolving dynamics. The findings demonstrate the capability of cooperative learning for decentralized redistribution and improving equitable service under uncertain and evolving dynamics.
发表机构
- North Carolina State University(北卡罗来纳州立大学)
- Bangladesh University of Engineering and Technology(孟加拉国工程技术大学)
机构由 AI 辅助整理,请以论文原文为准。