arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

使用多智能体强化学习解决空中走廊降级监视下的冲突

Conflict Resolution under Degraded Surveillance in Air Corridors Using Multi-Agent Reinforcement Learning

Esrat Farhana Dulia, Syed Arbab Mohd Shihab, Caleb Adams, Ruben Del Rosario

arXiv 2607.20547首次发表:更新:

发表机构

College of Aeronautics and Engineering; Center for Advanced Air Mobility(航空与工程学院; 先进空中交通中心)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究针对空中走廊降级监视下的冲突解决问题,开发基于深度Q网络的多智能体强化学习框架,为两类飞机训练策略,经模拟评估策略效果,揭示安全与走廊容量权衡,支持相关冲突解决策略评估。

AI 中文摘要

安全的先进空中交通行动要求飞机在监视信息嘈杂、延迟、不完整或暂时不可用时保持间隔。本研究开发了一种基于深度Q网络的多智能体强化学习框架,用于在结构化三维走廊内运行的异构小型无人机和电动垂直起降飞机之间的分散冲突解决。使用局部观测和14个动作空间为两类飞机训练单独的策略。模拟纳入了特定于飞机的动力学、能源使用、走廊约束、观测噪声、通信延迟、信息丢失、风干扰、执行器不确定性和模型不确定性。在90种交通密度和最小间隔阈值组合上评估训练后的策略。间隔损失频率和持续时间通常随交通密度和间隔要求增加,不过大多数事件在1秒内得到解决。在安全条件下,智能体约79%的时间保持运动。冲突期间,转向占动作的33%,其次是保持运动占29%,速度控制占25%,垂直机动占13%。六种帕累托最优配置揭示了安全与走廊容量之间的权衡。该框架支持在降级监视条件下对更安全的先进空中交通冲突解决策略进行基于模拟的评估。

英文摘要

Safe Advanced Air Mobility operations require aircraft to maintain separation when surveillance information is noisy, delayed, incomplete, or temporarily unavailable. This study develops a Deep Q-Network-based Multi-Agent Reinforcement Learning framework for decentralized conflict resolution among heterogeneous small unmanned aerial vehicles and electric vertical takeoff and landing aircraft operating within a structured three-dimensional corridor. Separate policies are trained for the two aircraft categories using local observations and a 14-action space that includes maintaining course, turning, vertical maneuvering, landing, and speed control. The simulation incorporates aircraft-specific dynamics, energy use, corridor constraints, observation noise, communication delay, information dropout, wind disturbance, actuator uncertainty, and model uncertainty. The trained policies are evaluated across 90 combinations of traffic density and minimum separation thresholds. Loss-of-separation frequency and duration generally increase with traffic density and separation requirements, although most events are resolved within 1s. Under safe conditions, agents maintain their motion approximately 79% of the time. During conflicts, turning accounts for 33% of actions, followed by maintaining motion at 29%, speed control at 25%, and vertical maneuvers at 13%. Six Pareto-optimal configurations reveal trade-offs between safety and corridor capacity. The framework supports the simulation-based evaluation of safer AAM conflict-resolution strategies under degraded surveillance conditions.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑