arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

BVR Sim:用于异构空战强化学习的开源高通量环境

BVR Sim: An Open and High-Throughput Environment for Heterogeneous Air-Combat Reinforcement Learning

Haocheng Sun, Mulai Tan

arXiv 2608.25419首次发表:更新:

发表机构

Beijing University of Posts and Telecommunications; Aviation Engineering School, Air Force Engineering University(北京邮电大学; 空军工程大学航空工程学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出开源异构空战强化学习环境 BVR Sim,支持多飞行器模型,具备高仿真效率,策略可跨平台迁移,且兼容主流多智能体强化学习框架。

AI 中文摘要

超视距(Beyond-visual-range,BVR)空战是一个极具挑战性的强化学习领域,其特征包括部分可观测性、长 horizon 决策、能量管理以及武器有限。本文提出 BVR Sim,这是一个开源的 Gymnasium 风格环境,专为异构空战强化学习设计。BVR Sim 支持多种 JSBSim 飞行器模型,包括 F-15、F-16、F/A-18 和 F-22,具备可配置的武器、传感器、控制器和对手。统一战术动作接口指定期望的航向、高度、速度和武器释放,位于飞行器特定的内环控制器之上,使策略可跨异构平台运行。该环境提供可互换的 Python 和加速 C++ 后端、面向实体的观测、组合式奖励、脚本化对手、回放与可视化,以及多智能体学习框架的适配器。在 0.4 秒决策间隔下,C++ 后端在 1 对 1 场景中达到每墙钟秒 104 个模拟秒,在 10 对 10 场景中仍保持实用。仅在 F-16 上训练的策略无需重新训练即可迁移至四种未见过的飞行器,通过飞行器特定控制器适配达到 45.5% 的平均胜率。MAPPO 和 HAPPO 实验进一步验证了其与标准多智能体强化学习流水线的端到端兼容性。

英文摘要

Beyond-visual-range (BVR) air combat is a challenging reinforcement-learning domain characterized by partial observability, long-horizon decision making, energy management, and limited weapons. We present BVR Sim, an open-source Gymnasium-style environment designed for heterogeneous air-combat reinforcement learning. BVR Sim supports multiple JSBSim aircraft models, including the F-15, F-16, F/A-18, and F-22, with configurable weapons, sensors, controllers, and opponents. A unified tactical action interface specifies desired heading, altitude, speed, and weapon release above aircraft-specific inner-loop controllers, enabling policies to operate across heterogeneous platforms. The environment provides interchangeable Python and accelerated C++ backends, entity-oriented observations, compositional rewards, scripted opponents, replay and visualization, and adapters for multi-agent learning frameworks. At a 0.4-s decision interval, the C++ backend achieves 104 simulated seconds per wall-clock second in 1-vs-1 and remains practical through 10-vs-10 scenarios. A policy trained only on the F-16 transfers without retraining to four unseen aircraft, reaching a 45.5% mean win rate with aircraft-specific controller adaptation. MAPPO and HAPPO experiments further verify end-to-end compatibility with standard multi-agent reinforcement-learning pipelines.

Comments9 pages, 4 figures, 4 tables. Code and reproducibility artifacts available at the project repository

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑