CosFly-VLA:一种用于无人机跟踪的空间感知视觉-语言-动作模型
CosFly-VLA: A Spatially Aware Vision-Language-Action Model for UAV Tracking
浏览论文内容
中文总结 AI 辅助
针对复杂城市环境中无人机动态目标跟踪问题,提出CosFly-VLA模型,通过结构化预测接口联合定位目标、估计可见性并生成飞行动作。经多阶段训练,该模型在开环误差和闭环成功率上有显著提升,实现从可见帧模仿到空间基础动作闭环控制的进展。
中文摘要 AI 辅助
动态目标跟踪对复杂城市环境中的无人机至关重要,现有视觉-语言-动作(VLA)策略在视线受阻时性能会下降。为此提出CosFly-VLA模型,它通过结构化预测接口联合定位目标、估计其可见性并生成连续飞行动作。训练时,在500k混合池上进行空间基础持续预训练,通过三阶段基于课程的监督微调、思维链训练及闭环强化学习。实验表明,相对于OpenVLA,CosFly-VLA-0.8B在可见测试和不可见测试中分别降低了开环平均位移误差,闭环优化提高了成功率。这些结果展示了从可见帧模仿到空间基础动作闭环控制的进展。
英文摘要
Dynamic target tracking is essential for Unmanned Aerial Vehicles (UAVs) operating in complex urban environments, where both the target and the camera viewpoint change continuously. Existing Vision-Language-Action (VLA) policies can track visible targets effectively, but their performance often degrades when buildings, vegetation, or roadside objects block the line of sight. During sustained occlusion, a policy may lose the target state, execute actions toward an incorrect region, and amplify this error through subsequent observations until re-acquisition becomes impossible. To this end, we present CosFly-VLA, a spatially aware VLA model that jointly grounds the target, estimates its visibility, and generates continuous flight actions through a structured prediction interface. To train this policy, we use a large-scale recipe over diverse data sources. Spatially Grounded Continued Pretraining (CPT) on a 500k mixed pool injects UAV-view depth, distance, and 3-D spatial reasoning. A three-stage Curriculum-based Supervised Fine-Tuning (SFT) process then specializes the tracker through multi-head warm-up followed by two-stage curriculum learning over natural and hard / long-occlusion data. Chain-of-Thought (CoT) training subsequently teaches recovery-oriented reasoning traces before structured answers. Finally, a closed-loop Reinforcement Learning (RL) stage optimizes tracking behavior with a multi-component reward covering stand-off tracking, grounding quality, collision avoidance, and task success. Relative to OpenVLA, CosFly-VLA-0.8B reduces open-loop Average Displacement Error (ADE) by 34.1% on seen-test and 35.3% on unseen-test. Closed-loop optimization improves Success Rate (SR) by 29.8% and 2.5%, respectively. These results demonstrate progress from visible-frame imitation toward spatially grounded action-closed-loop control, evaluated under a shared oracle state history.
发表机构
- Autel Robotics(大疆创新科技有限公司)
- Northeast Normal University(东北师范大学)
- Southern University of Science and Technology(南方科技大学)
- Peking University(北京大学)
- University of Hong Kong(香港大学)
机构由 AI 辅助整理,请以论文原文为准。