发表机构
University of São Paulo (USP)(圣保罗大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对空中机器人机载视觉语言导航的挑战,提出VLN on the Fly模块化栈,分离接地、规划与控制,在15次飞行中13次成功,平均误差5.72厘米。
AI 中文摘要
在空中机器人上完全机载地运行视觉语言导航是困难的,因为接地(grounding)、规划和控制必须共享有限的计算资源,且单级错误在飞行中难以隔离。端到端的空中策略将这些阶段融合到一个网络中,放弃了模块化栈所保留的可观测性和安全检查。我们提出VLN on the Fly,一个机载栈,将接地、规划和控制保持为分离的、可检查的阶段。一个量化VLM将指令接地到一个粗略的图像单元,深度将其提升到3D目标,一个快速的B样条规划器返回可行的轨迹,一个预训练的强化学习策略将其跟踪到跨四旋翼的电机命令。在受控室内体积中对三个日常指称物进行的15次机载飞行中,该栈在13/15次试验中到达目标,平均目标误差为5.72厘米,平均GPU利用率为39.3%。在另外6次杂乱环境试验中,该栈在机载感知门控下跟踪无碰撞轨迹。
英文摘要
Running vision-language navigation fully onboard an aerial robot is hard, since grounding, planning, and control must share limited compute and a single-stage error is difficult to isolate in flight. End-to-end aerial policies fuse these stages into one network, giving up the observability and safety checks a modular stack keeps available. We propose VLN on the Fly, an onboard stack that keeps grounding, planning, and control as separate, inspectable stages. A quantized VLM grounds an instruction to a coarse image cell, depth lifts it to a 3D goal, a fast B-spline planner returns a feasible trajectory, and a pretrained reinforcement learning policy tracks it to motor commands across quadrotors. Across 15 onboard flights over three everyday referents in a controlled indoor volume, the stack reaches the target in 13 of 15 trials with 5.72 cm mean goal error and 39.3% average GPU utilization. In 6 additional cluttered-environment trials, the stack tracks collision-free trajectories under onboard perception gating.