arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

LightVLN:具有紧凑记忆与历史引导局部聚合的高效空中视觉语言导航

LightVLN: Efficient Aerial Vision-and-Language Navigation with Compact Memory and History-Guided Local Aggregation

Yiming Zhao, Tianshun Li, Jingle He, Ruonan Chai, Xinhu Zheng

arXiv 2610.05024首次发表:更新:

发表机构

The Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州))

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

LightVLN提出一种轻量级空中视觉语言导航框架,通过紧凑记忆与历史引导局部聚合降低计算开销,在OpenFly和AerialVLN-S数据集上超越7B基线,并实现机载实时部署。

AI 中文摘要

空中视觉语言导航(VLN)使无人机能够在复杂的三维环境中,根据视觉观察执行长时程的自然语言指令。然而,现有的空中VLN模型通常依赖大规模的视觉-语言骨干网络和密集的视觉历史,带来了巨大的计算和内存开销,阻碍了机载部署。我们提出了LightVLN,一个轻量级的历史感知空中VLN框架,它结合了紧凑的0.5B语言骨干网络以及历史和当前观察的紧凑表示。LightVLN利用策略已计算的视觉特征,将每个历史帧压缩为单个令牌。它进一步引入历史和指令条件化的局部聚合,将当前观察从256个视觉令牌减少到32个,同时保留与导航相关的空间信息。在最多16个历史帧的情况下,策略最多使用48个观察派生令牌。在公开的OpenFly数据集上,LightVLN在Test-Seen和Test-Unseen上的成功率(SR)分别达到50.93%和36.14%,在大多数报告的指标上优于所评估的7B语言骨干基线。它还在AerialVLN-S Val-Seen上实现了25.83%的SR。在一个重建的未见校园中,我们将LightVLN部署在DJI M350 RTK上,并外接Jetson Orin NX 16 GB,进行闭环机载计算的真实到模拟硬件在环(HIL)评估,实现了14.61 Hz的模型推理和11.13 Hz的端到端决策更新。这些结果证明了LightVLN在航空导航中的有效性和效率。

英文摘要

Aerial vision-and-language navigation (VLN) enables unmanned aerial vehicles to execute long-horizon natural-language instructions from visual observations in complex three-dimensional environments. However, recent aerial VLN models often rely on large-scale vision-language backbones and dense visual histories, imposing substantial computation and memory costs that hinder onboard deployment. We propose LightVLN, a lightweight history-aware aerial VLN framework that combines a compact 0.5B language backbone with compact representations of both historical and current observations. LightVLN compresses each historical frame into a single token using visual features already computed by the policy. It further introduces history- and instruction-conditioned local aggregation to reduce the current observation from 256 to 32 visual tokens while preserving navigation-relevant spatial information. With up to 16 historical frames, the policy uses at most 48 observation-derived tokens. On the public OpenFly dataset, LightVLN achieves 50.93% Test-Seen and 36.14% Test-Unseen success rates (SR), outperforming the evaluated 7B language-backbone baselines on most reported metrics. It also achieves 25.83% SR on AerialVLN-S Val-Seen. In a reconstructed unseen campus, we deploy LightVLN on a DJI M350 RTK with an external Jetson Orin NX 16 GB for closed-loop onboard-compute real-to-sim hardware-in-the-loop (HIL) evaluation, achieving 14.61 Hz model inference and 11.13 Hz end-to-end decision updates. These results demonstrate the effectiveness and efficiency of LightVLN for aerial navigation.

Comments8 pages, 5 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑