ForeFly:用于空中视觉语言导航的双视界世界动作模型
ForeFly: A Dual-Horizon World Action Model for Aerial Vision-Language Navigation
浏览论文内容
中文总结 AI 辅助
提出ForeFly双视界世界动作模型,通过近端与路径关键未来预测及FGAR细化,提升空中视觉语言导航在复杂环境中的指令遵循性能。
中文摘要 AI 辅助
空中视觉语言导航(AVLN)要求无人机在复杂的三维环境中沿长轨迹保持可靠的指令遵循。然而,现有的AVLN方法主要反应式或仅限于单视界预测,忽略了不同时间视界上的互补未来线索。为解决此局限,我们提出ForeFly,一种双视界潜在世界动作模型,同时预测用于局部连续性的近端未来和用于远距离引导的自适应路径关键未来。视界特定的前瞻查询以近期和路径关键的视觉记忆为初始,为未来预测提供历史感知上下文。为利用它们在动作生成中的不同角色,我们引入前瞻引导的动作细化(FGAR),其不对称地利用近端前瞻进行局部动作增强,利用路径关键前瞻进行特征级修正和路径级引导。在TravelUAV和UAV-ON基准上的实验表明,ForeFly在已知和未知设置中均持续优于强基线,验证了双视界前瞻和FGAR学习的有效性。代码可在以下网址获取:this https URL
英文摘要
Aerial Vision-Language Navigation (AVLN) requires UAVs to maintain reliable instruction following over long trajectories in complex 3D environments. However, existing AVLN approaches are predominantly reactive or limited to single-horizon prediction, overlooking complementary future cues across different temporal horizons. To address this limitation, we propose ForeFly, a dual-horizon latent world action model that predicts both a proximal future for local continuity and an adaptive route-critical future for long-range guidance. Horizon-specific foresight queries are primed with recent and route-critical visual memories, providing history-aware context for future prediction. To exploit their distinct roles in action generation, we introduce Foresight-Guided Action Refinement (FGAR), which asymmetrically exploits proximal foresight for local action enhancement and route-critical foresight for feature-wise correction and route-level guidance. Experiments on the TravelUAV and UAV-ON benchmarks show that ForeFly consistently outperforms strong baselines across seen and unseen settings, validating the effectiveness of dual-horizon foresight and FGAR learning. The code is available at: https://github.com/kunhuiW/ForeFly
发表机构
- Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所)
- University of Chinese Academy of Sciences(中国科学院大学)
- Peng Cheng Laboratory(鹏城实验室)
- Duke Kunshan University(昆山杜克大学)
机构由 AI 辅助整理,请以论文原文为准。