发表机构
University of Haifa(海法大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
提出混合保真度训练方案,用离线残差校正低保真模拟器训练共享去中心化策略,避免高保真强化学习,在多种集群任务中性能接近高保真微调策略且计算成本大幅降低。
AI 中文摘要
直接在低保真(LF)刚体物理中训练多智能体无人机集群策略虽然准确,但计算成本高昂。该成本随团队规模扩大而急剧增加,因为每增加一个智能体都会使接触解析复杂度成倍增长,并显著提高模拟中的碰撞率。为解决此问题,我们提出了一种混合保真度训练方案,完全消除了对高保真(HF)强化学习的需求。一个共享的、去中心化的策略在完全可微的、基于JAX的低保真点质量模拟器中进行优化。该模拟器通过一个小的、每智能体袋装残差集成进行校正,该集成仅需在HF模拟器中使用短校准飞行离线拟合一次。由于校准仅需一架隔离的无人机,数据收集预算不会随团队规模而增加。参考轨迹通过滚动执行现有的仅低保真策略生成,并由零训练PD控制器在HF模拟器中跟踪。在四个协作无人机任务和3至18的团队规模评估中,残差校正策略在所有组合中均优于未校正的低保真基线,并在测试的24种组合中的22种中优于从头训练的高保真策略。它与高保真微调策略的差距随团队规模增大而稳步缩小。最终,所提方法在最大团队规模下以极低的计算成本实现了近乎等效的性能,完全避免了高保真训练中典型的高碰撞率。
英文摘要
Training multi-agent drone-swarm policies directly in high-fidelity (HF) rigid-body physics is accurate but computationally expensive. This cost scales poorly with team size, as each additional agent multiplies contact-resolution complexity and sharply raises the in-simulation crash rate. To address this, we propose a mixed-fidelity training scheme that eliminates HF reinforcement learning entirely. A single shared, decentralized policy is optimized inside a fully-differentiable, JAX-native low-fidelity (LF) point-mass simulator. The simulator is corrected by a small, per-agent bagged residual ensemble fit once, offline, using short calibration flights in the HF simulator. Because calibration requires only one isolated drone, the data collection budget does not compound with team size. Reference trajectories are generated by rolling out an existing LF-only policy and tracked in the HF simulator by a zero-training PD controller. Evaluated across four cooperative drone tasks and team sizes from 3 to 18, the residual-corrected policy outperforms an uncorrected LF baseline in all combinations, and a from-scratch HF policy in 22 of 24 combinations tested. It trails an HF-finetuned policy by a margin that narrows steadily with team size. Ultimately, the proposed method achieves near-equivalent performance at the largest team sizes at a fraction of the computational cost, completely avoiding the high crash rates typical of HF training.
Comments8 pages, 5 figures