arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.15004cs.RO

CosFly-VLA:一种用于无人机跟踪的空间感知视觉-语言-动作模型

CosFly-VLA: A Spatially Aware Vision-Language-Action Model for UAV Tracking

Ruilong Ren, Songsheng Cheng, Yunpeng Zhou, Hanxuan Chen, Xiangyue Wang, Tianle Zeng, Shuai Yuan, Binbo Li, Hanzhong Guo, Ji Pei, Da Zhang, Kangli Wang

首次发表
浏览论文内容

中文总结 AI 辅助

针对复杂城市环境中无人机动态目标跟踪问题,提出CosFly-VLA模型,通过结构化预测接口联合定位目标、估计可见性并生成飞行动作。经多阶段训练,该模型在开环误差和闭环成功率上有显著提升,实现从可见帧模仿到空间基础动作闭环控制的进展。

中文摘要 AI 辅助

动态目标跟踪对复杂城市环境中的无人机至关重要,现有视觉-语言-动作(VLA)策略在视线受阻时性能会下降。为此提出CosFly-VLA模型,它通过结构化预测接口联合定位目标、估计其可见性并生成连续飞行动作。训练时,在500k混合池上进行空间基础持续预训练,通过三阶段基于课程的监督微调、思维链训练及闭环强化学习。实验表明,相对于OpenVLA,CosFly-VLA-0.8B在可见测试和不可见测试中分别降低了开环平均位移误差,闭环优化提高了成功率。这些结果展示了从可见帧模仿到空间基础动作闭环控制的进展。

英文摘要

Dynamic target tracking is essential for Unmanned Aerial Vehicles (UAVs) operating in complex urban environments, where both the target and the camera viewpoint change continuously. Existing Vision-Language-Action (VLA) policies can track visible targets effectively, but their performance often degrades when buildings, vegetation, or roadside objects block the line of sight. During sustained occlusion, a policy may lose the target state, execute actions toward an incorrect region, and amplify this error through subsequent observations until re-acquisition becomes impossible. To this end, we present CosFly-VLA, a spatially aware VLA model that jointly grounds the target, estimates its visibility, and generates continuous flight actions through a structured prediction interface. To train this policy, we use a large-scale recipe over diverse data sources. Spatially Grounded Continued Pretraining (CPT) on a 500k mixed pool injects UAV-view depth, distance, and 3-D spatial reasoning. A three-stage Curriculum-based Supervised Fine-Tuning (SFT) process then specializes the tracker through multi-head warm-up followed by two-stage curriculum learning over natural and hard / long-occlusion data. Chain-of-Thought (CoT) training subsequently teaches recovery-oriented reasoning traces before structured answers. Finally, a closed-loop Reinforcement Learning (RL) stage optimizes tracking behavior with a multi-component reward covering stand-off tracking, grounding quality, collision avoidance, and task success. Relative to OpenVLA, CosFly-VLA-0.8B reduces open-loop Average Displacement Error (ADE) by 34.1% on seen-test and 35.3% on unseen-test. Closed-loop optimization improves Success Rate (SR) by 29.8% and 2.5%, respectively. These results demonstrate progress from visible-frame imitation toward spatially grounded action-closed-loop control, evaluated under a shared oracle state history.

发表机构

  • Autel Robotics(大疆创新科技有限公司)
  • Northeast Normal University(东北师范大学)
  • Southern University of Science and Technology(南方科技大学)
  • Peking University(北京大学)
  • University of Hong Kong(香港大学)

机构由 AI 辅助整理,请以论文原文为准。

↑