arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

VLM-MPPI:在行为多样轨迹中锚定自然语言用于空中导航

VLM-MPPI: Grounding Natural Language in Behaviorally Diverse Trajectories for Aerial Navigation

Hanbing Zhang, Fangguo Zhao, Zerui Li, Xin Guan, Peng Cheng, Shuo Li

arXiv 2609.18451首次发表:更新:

发表机构

Zhejiang University; Australian Institute for Machine Learning, Adelaide University(浙江大学; 阿德莱德大学澳大利亚机器学习研究所)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

提出VLM-MPPI分层无人机导航框架,通过六个行为条件MPPI规划器生成多样轨迹,并利用预训练VLM将自然语言意图映射为视觉动作选择,在仿真和真实实验中实现100%任务成功率。

AI 中文摘要

我们提出了一种分层无人机导航框架,该框架在杂乱室内环境中将自然语言意图与动态可行的飞行行为对齐。为弥合抽象语义与低级控制之间的鸿沟,我们采用了一个由六个行为条件模型预测路径积分(MPPI)规划器组成的并行集成。关键在于,通过设计特定模式的引导成本和采样偏差,我们诱导出收敛于独特行为均值的不同轨迹模式,从而产生一组紧凑的、有意多样的候选轨迹,而非仅仅是随机变化。我们将这些3D候选轨迹投影到机载第一人称视角RGB流上,将语言锚定转化为视觉动作选择问题。一个预训练的视觉-语言模型(VLM)异步地根据叠加的FPV图像和自然语言提示选择候选索引,而MPPI以20Hz的频率重新规划,基于PID的低级控制器跟踪所选轨迹。我们在NVIDIA Isaac Sim中以及配备LiDAR和RGB传感的真实四旋翼平台上实现了完整流程。在仿真和真实飞行中的实验表明,行为多样性具有语义意义,尽管存在VLM延迟,语言对齐依然稳健,且在所有模式下飞行安全且可重复,在我们评估的场景中实现了100%的任务成功率。

英文摘要

We present a hierarchical UAV navigation framework that aligns natural-language intent with dynamically feasible flight behaviors in cluttered indoor environments. To bridge the gap between abstract semantics and low-level control, we employ a parallelized ensemble of six behavior-conditioned Model Predictive Path Integral (MPPI) planners. Crucially, by designing mode-specific guiding costs and sampling biases, we induce distinct trajectory modes that converge to unique behavioral means, yielding a compact set of intentionally diverse candidates rather than mere stochastic variations. We project these 3D candidates onto the onboard first-person-view RGB stream, turning language grounding into a visual action selection problem. A pretrained vision--language model (VLM) asynchronously selects the candidate index given the overlaid FPV image and a natural-language prompt, while MPPI replans at 20Hz and a PID-based low-level controller tracks the selected trajectory. We implement the full pipeline in NVIDIA Isaac Sim and on a real-world quadrotor platform equipped with LiDAR and RGB sensing. Experiments in both simulation and real-world flights show semantically meaningful behavior diversity, robust language alignment despite VLM latency, and safe, repeatable flight across all modes, achieving 100% task success in our evaluated scenarios.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑