arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.22925cs.RO

“亲爱的LLaVA,请驾驶”:一种用于闭环机器人控制的深度感知视觉语言智能体

"Dear LLaVA, Please Drive": A Depth-Aware Vision-Language Agent for Closed-Loop Robotic Control

Sebastian Berger, Katharina Winter, Fabian B. Flohr

首次发表
浏览论文内容

中文总结 AI 辅助

该研究提出一种基于命令式学习和LoRA的深度感知视觉语言智能体,仅用立体深度观测即可生成无碰撞路径,在更新不到1%参数的情况下实现高效闭环导航。

中文摘要 AI 辅助

视觉语言模型(VLMs)为基于推理的移动导航提供了引人注目的基础,通过大规模预训练提供了丰富的上下文理解和强大的泛化能力。大多数现有的导航框架依赖于模仿学习,因此需要大量带标签的轨迹数据,限制了其可扩展性和鲁棒性。在这项工作中,我们提出了一种参数高效的方法,使用命令式学习(Imperative Learning)范式对预训练的VLM进行微调,用于自主导航。通过针对可微分的几何代价场而非带标签轨迹进行优化,我们的模型仅从立体深度观测中学习生成无碰撞路径。我们引入了一个统一的端到端导航流水线,用于自然语言驱动的机器人控制。该系统利用共享的VLM骨干网络和任务特定的低秩适应(LoRA)模块,有效地弥合了从语义目标选择到低级轨迹规划之间的差距。我们的方法在未见环境中取得了具有竞争力的路径长度加权成功率(SPL),同时仅更新了模型总参数的不到1%。定性的真实世界实验验证了无需在真实世界数据上微调即可实现仿真到现实的泛化和稳定的路径规划。这些结果突显了一种在移动机器人上部署基于VLM的智能体的实用方法,使得无需大规模带标签轨迹数据的苛刻要求即可实现高级语义导航。

英文摘要

Vision-language models (VLMs) provide a compelling foundation for reasoning-driven mobile navigation, offering rich contextual understanding and strong generalization from large-scale pretraining. Most existing navigation frameworks rely on imitation learning and therefore require substantial labeled trajectory data, limiting their scalability and robustness. In this work, we propose a parameter-efficient approach to fine-tune a pretrained VLM for autonomous navigation using an Imperative Learning paradigm. By optimizing against differentiable geometric cost fields rather than labeled trajectories, our model learns to generate collision-free paths exclusively from stereoscopic depth observations. We introduce a unified end-to-end navigation pipeline for natural-language-driven robotic control. This system leverages a shared VLM backbone with task-specific Low-Rank Adaptation (LoRA) modules, effectively bridging the gap from semantic target selection to low-level trajectory planning. Our approach achieves competitive Success weighted by Path Length (SPL) in unseen environments while updating less than 1% of the model's total parameters. Qualitative real-world experiments validate sim-to-real generalization and stable path planning without fine-tuning on real-world data. These results highlight a practical approach for deploying VLM-based agents on mobile robots, enabling high-level semantic navigation without the prohibitive requirement for large-scale, labeled trajectory data.

发表机构

  • Munich University of Applied Sciences(慕尼黑应用科学大学)

机构由 AI 辅助整理,请以论文原文为准。

↑