AI 中文总结
RecoverFly是面向端到端无人机视觉-语言-动作策略的故障感知强化学习后训练框架,在TravelUAV基准上实现了优于AerialVLA的性能,提升了导航成功率。
AI 中文摘要
无人机视觉语言导航(UAV-VLN)要求智能体在复杂环境中将视觉观测和语言指令转化为可靠的飞行动作。尽管近期的端到端无人机视觉-语言-动作(UAV-VLA)策略降低了对单独设计的感知、规划和控制模块的依赖,但它们的行为克隆目标为交互式闭环执行提供的校正监督有限。强化学习(RL)是一种有前景的解决方案,但其有效性受限于样本利用效率低、场景分布长尾性以及优化过程中的策略分布偏移问题。为此,我们提出RecoverFly,这是一种面向端到端UAV-VLA策略的故障感知强化学习后训练框架。具体而言,RecoverFly将标记级强化学习用于语法约束的自回归UAV动作的稳定优化,重新处理未解决的故障案例以增强校正学习和样本利用,并结合两阶段长尾场景课程与参考策略正则化,在保留已习得能力的同时提升场景适应性。在TravelUAV基准上的实验表明,RecoverFly在可见、未见地图和未见对象分割上均取得最佳性能。此外,与AerialVLA初始化相比,在总部署预算约为训练集规模30%的情况下,RecoverFly将成功率提升了3.12至8.37个百分点,验证了其有效性、鲁棒性和泛化能力。
英文摘要
Unmanned aerial vehicle vision-language navigation (UAV-VLN) requires agents to translate visual observations and language instructions into reliable flight actions in complex environments. Although recent end-to-end UAV vision-language-action (UAV-VLA) policies reduce reliance on separately designed perception, planning, and control modules, their behavior-cloning objectives provide limited supervision for errors arising during closed-loop execution. Reinforcement learning (RL) offers a promising solution, while limited interaction budgets, long-tailed scene distributions, and policy drift complicate sustained learning. To this end, we propose RefineFly, a failure-aware RL post-training framework for end-to-end UAV-VLA policies. Built on token-level proximal policy optimization (PPO), RefineFly maintains dynamic failure memory within a two-stage scene curriculum to sustain learning from unresolved tasks as training shifts from empirical to balanced scene sampling. Throughout this process, stage-specific reference regularization constrains policy deviation. Experiments on the TravelUAV benchmark demonstrate that RefineFly outperforms all comparison methods across all three splits, improving success rate over AerialVLA by 3.12 to 8.37 percentage points with a total rollout budget of about 30\% of the training-set size. Moreover, ablations reveal that the benefits of sustained failure learning vary across evaluation splits, while the two-stage scene curriculum improves overall performance, highlighting the importance of training distributions when learning from failures.