发表机构
Appsofa LLC(Appsofa有限责任公司)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究提出UC-PSRO方法,针对通信降级环境中对抗自适应对手的无人机集群,经实验发现其通信失活课程可提升任务鲁棒性,但效用条件化与PSRO自博弈会减慢收敛速度。
AI 中文摘要
本研究受美国空军公开SBIR提案的启发(但并非源自该提案),旨在为通信降级环境中对抗自适应红方对手的蓝方无人机集群生成博弈论优化的行动方案(COA)。我们提出UC-PSRO(Utility-Conditioned Policy-Space Response Oracles with a Communication-Dropout Curriculum,即带通信失活课程的效用条件策略空间响应神谕),该方法融合了三种机制:(i)PSRO自博弈,使蓝方与红方策略互为近似最优响应进行训练,而非一方对抗固定脚本化对手;(ii)基于指挥官意图权重向量的FiLM条件化蓝方策略,该向量在训练时从狄利克雷分布中采样,使训练好的策略在执行时无需重新训练即可重新导向;(iii)训练期间采用课程退火通信图边失活,使集群学习去中心化的点对点 fallback,而非依赖全连接。我们针对该提案海上场景的合成非机密替代场景进行评估,使用5个随机种子,蓝方智能体数量N=25,并扩展至N=200的可扩展性扫描。研究发现存在真实的权衡,而非统一的胜负:仅通信失活课程在所有学习方法中提供最强、最鲁棒的任务完成率,且随通信拒止程度增加反而提升(失活率从0升至0.75时,成功率从35%升至62%);添加效用条件化和PSRO自博弈会在固定预算内显著减慢收敛速度,且我们发现自博弈相比固定对手策略无可靠可利用性优势,两者统计上无差异,差距极小接近零。我们如实报告这一尚未被鲁棒性优势抵消的收敛成本,而非夸大某一方法为主导,并提供完全向量化的开放环境,在单张消费级GPU上,N=200个智能体每步推理耗时仅为个位数毫秒。
英文摘要
We study generating game-theoretically optimized Courses of Action (COAs) for a Blue UAS swarm against an adaptive Red adversary in a communication-degraded environment, motivated by (but not derived from) a public U.S. Air Force SBIR solicitation. We propose UC-PSRO (Utility-Conditioned Policy-Space Response Oracles with a Communication-Dropout Curriculum), combining three mechanisms: (i) PSRO self-play, so Blue and Red policies train as approximate best responses to each other rather than one side against a fixed scripted opponent; (ii) FiLM conditioning of the Blue policy on a Commander's-Intent weight vector, sampled from a Dirichlet distribution during training, so one trained policy is re-steerable at execution time without retraining; and (iii) a curriculum annealing communication-graph edge dropout during training, so the swarm learns decentralized, peer-to-peer fallback instead of depending on full connectivity. We evaluate on a synthetic, unclassified stand-in for the solicitation's maritime scenario, with 5 seeds at N=25 Blue agents and a scalability sweep to N=200. We find a genuine trade-off, not a uniform win: the communication-dropout curriculum alone gives the strongest, most robust mission-completion rates of any learned method, improving counter-intuitively as denial increases (35% to 62% success as dropout rises from 0 to 0.75); adding utility-conditioning and PSRO self-play substantially slows convergence within a fixed budget, and we find no reliable exploitability advantage for self-play over a fixed-opponent policy, both statistically indistinguishable from a small, near-zero gap. We report this honestly as a convergence cost not yet offset by a demonstrated robustness benefit, rather than overstating one method as dominant, and provide a fully vectorized, open environment training at N=200 agents in single-digit milliseconds per step on a single consumer GPU.