arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

DreamFly:面向空中视觉语言导航的因果记忆与后退时域扩散规划

DreamFly: Causal Memory and Receding-Horizon Diffusion Planning for Aerial Vision-Language Navigation

Yan Deng, Fei Xu

arXiv 2608.12308首次发表:更新:

发表机构

School of Electronic Information Engineering, Xi’an Technological University; School of Computer Science and Engineering, Xi’an Technological University(西安工业大学电子信息工程学院; 西安工业大学计算机科学与工程学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

DreamFly是基于Dream-VLA的扩散式空中VLN框架,通过因果记忆、后退时域扩散规划和LiteStop解耦终止,在OpenFly基准上显著优于对比方法。

AI 中文摘要

空中视觉语言导航(VLN)要求具身智能体在部分可观测环境下,整合随时间变化的视觉证据、规划未来动作并判断是否到达导航目标。尽管近期的VLA模型提供了从感知到动作的有前景范式,但将其适配到空中导航仍面临挑战,具体包括历史上下文有限、规划时域较短以及隐式终止不可靠。为解决这些问题,我们提出了基于Dream-VLA的扩散式空中VLN框架DreamFly。DreamFly引入了因果对齐的历史记忆,仅使用当前决策步骤之前的观测来增强当前视觉表示,从而在无未来信息泄露的情况下实现时间推理。我们进一步将导航问题建模为后退时域扩散规划,其中策略预测K步动作块,但仅执行第一步动作后重新规划。这种“规划K步、执行一步”的策略将未来动作作为辅助规划目标,同时保留闭环视觉反馈。最后,LiteStop在初始全掩码状态下直接从动作逻辑估计停止概率,将显式终止与动作生成解耦。在OpenFly基准上的实验表明,DreamFly在可见和不可见环境中均取得了一致的性能提升:在测试可见/测试不可见划分上,分别达到32.04%/29.46%的成功率(SR)和28.22%/23.54%的路径规划成功率(SPL),在两项指标上均优于所有对比方法,同时实现了最低的导航误差。这些结果证明,对历史上下文、未来动作结构和显式终止进行联合建模,对空中VLN具有有效性。

英文摘要

Aerial vision-language navigation (VLN) requires an embodied agent to integrate visual evidence over time, plan future actions, and determine when it has reached a navigation goal under partial observability. Although recent VLA models offer a promising perception-to-action paradigm, adapting them to aerial navigation remains challenging due to limited historical context, short planning horizons, and unreliable implicit termination. To address these challenges, we propose DreamFly, a diffusion-based aerial VLN framework built on Dream-VLA. DreamFly introduces a causally aligned historical memory that augments the current visual representation using only observations preceding the current decision step, enabling temporal reasoning without future information leakage. We further formulate navigation as receding-horizon diffusion planning, where the policy predicts a $K$-step action chunk but executes only the first action before replanning. This plan-$K$, execute-one strategy uses future actions as auxiliary planning targets while preserving closed-loop visual feedback. Finally, LiteStop estimates the stop probability directly from action logits at the initial all-mask state, decoupling explicit termination from action generation. Experiments on the OpenFly benchmark demonstrate consistent improvements in seen and unseen environments. DreamFly achieves 32.04%/29.46% SR and 28.22%/23.54% SPL on the test-seen/test-unseen splits, respectively, outperforming all compared methods on both metrics while attaining the lowest navigation error. These results demonstrate the effectiveness of jointly modeling historical context, future action structure, and explicit termination for aerial VLN.

Comments24 pages, 6 figures, 3 tables

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑