arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

VLN on the Fly:面向空中机器人的机载视觉语言导航栈

VLN on the Fly: An Onboard Vision-Language Navigation Stack for Aerial Robots

Marco S. Tayar, Felipe Tommaselli, Gianluca Capezutto, Pedro Antonio Rabelo Saraiva, Pedro H. V. de Freitas, Lucas Kido, Guilherme Sonego, Ricardo V. Godoy, Marcelo Becker

arXiv 2609.20191首次发表:更新:

发表机构

University of São Paulo (USP)(圣保罗大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对空中机器人机载视觉语言导航的挑战,提出VLN on the Fly模块化栈,分离接地、规划与控制,在15次飞行中13次成功,平均误差5.72厘米。

AI 中文摘要

在空中机器人上完全机载地运行视觉语言导航是困难的,因为接地(grounding)、规划和控制必须共享有限的计算资源,且单级错误在飞行中难以隔离。端到端的空中策略将这些阶段融合到一个网络中,放弃了模块化栈所保留的可观测性和安全检查。我们提出VLN on the Fly,一个机载栈,将接地、规划和控制保持为分离的、可检查的阶段。一个量化VLM将指令接地到一个粗略的图像单元,深度将其提升到3D目标,一个快速的B样条规划器返回可行的轨迹,一个预训练的强化学习策略将其跟踪到跨四旋翼的电机命令。在受控室内体积中对三个日常指称物进行的15次机载飞行中,该栈在13/15次试验中到达目标,平均目标误差为5.72厘米,平均GPU利用率为39.3%。在另外6次杂乱环境试验中,该栈在机载感知门控下跟踪无碰撞轨迹。

英文摘要

Running vision-language navigation fully onboard an aerial robot is hard, since grounding, planning, and control must share limited compute and a single-stage error is difficult to isolate in flight. End-to-end aerial policies fuse these stages into one network, giving up the observability and safety checks a modular stack keeps available. We propose VLN on the Fly, an onboard stack that keeps grounding, planning, and control as separate, inspectable stages. A quantized VLM grounds an instruction to a coarse image cell, depth lifts it to a 3D goal, a fast B-spline planner returns a feasible trajectory, and a pretrained reinforcement learning policy tracks it to motor commands across quadrotors. Across 15 onboard flights over three everyday referents in a controlled indoor volume, the stack reaches the target in 13 of 15 trials with 5.72 cm mean goal error and 39.3% average GPU utilization. In 6 additional cluttered-environment trials, the stack tracks collision-free trajectories under onboard perception gating.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑