arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2603.10528cs.LGcs.AI

无人机多智能体强化学习:用于时间敏感和动态医疗物资配送

UAV-MARL: Multi-Agent Reinforcement Learning for Time-Critical and Dynamic Medical Supply Delivery

Islam Guven, Mehmet Parlak

更新

AI总结:

本文提出基于多智能体强化学习的无人机协调框架,用于在动态医疗物资配送中实现高效任务优先和资源分配。

AI中文摘要:

无人机(UAVs)日益被用于支持时间敏感的医疗物资配送,为紧急情况和资源短缺期间提供快速且灵活的物流支持。然而,有效部署无人机舰队需要能够优先考虑医疗需求、分配有限的空中资源并在不确定的运营条件下适应配送计划的协调机制。本文提出了一种多智能体强化学习(MARL)框架,用于协调在随机医疗配送场景中无人机舰队,其中请求在紧迫性、位置和配送截止时间上各不相同。该问题被建模为部分可观测马尔可夫决策过程(POMDP),其中无人机智能体保持对医疗配送需求的意识,同时由于通信和定位限制,对其他智能体的可见性有限。所提出的框架采用近端策略优化(PPO)作为主要学习算法,并评估了几种变体,包括异步扩展、经典演员-批评者方法以及架构修改,以分析可扩展性和性能权衡。该模型使用来自选定诊所和医院的现实世界地理数据,提取自OpenStreetMap数据集进行评估。该框架提供了一个决策支持层,优先处理医疗任务,实时重新分配无人机资源,并协助医疗人员管理紧急物流。实验结果表明,经典PPO在协调性能上优于异步和顺序学习策略,突显了强化学习在适应性和可扩展的无人机辅助医疗物流中的潜力。

英文摘要:

Unmanned aerial vehicles (UAVs) are increasingly used to support time-critical medical supply delivery, providing rapid and flexible logistics during emergencies and resource shortages. However, effective deployment of UAV fleets requires coordination mechanisms capable of prioritizing medical requests, allocating limited aerial resources, and adapting delivery schedules under uncertain operational conditions. This paper presents a multi-agent reinforcement learning (MARL) framework for coordinating UAV fleets in stochastic medical delivery scenarios where requests vary in urgency, location, and delivery deadlines. The problem is formulated as a partially observable Markov decision process (POMDP) in which UAV agents maintain awareness of medical delivery demands while having limited visibility of other agents due to communication and localization constraints. The proposed framework employs Proximal Policy Optimization (PPO) as the primary learning algorithm and evaluates several variants, including asynchronous extensions, classical actor--critic methods, and architectural modifications to analyze scalability and performance trade-offs. The model is evaluated using real-world geographic data from selected clinics and hospitals extracted from the OpenStreetMap dataset. The framework provides a decision-support layer that prioritizes medical tasks, reallocates UAV resources in real time, and assists healthcare personnel in managing urgent logistics. Experimental results show that classical PPO achieves superior coordination performance compared to asynchronous and sequential learning strategies, highlighting the potential of reinforcement learning for adaptive and scalable UAV-assisted healthcare logistics.

补充信息

↑