arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.14297cs.RO

多任务视觉感知网络与LLM条件化用于自主导航

Multi-Task Visual Perception Network with LLM Conditioning for Autonomous Navigation

  • Indian Institute of Technology Kanpur(印度理工学院坎普尔分校)

机构由 AI 辅助整理,请以论文原文为准。

Praveen Kumar, K. R. Guruprasad, Tushar Sandhan

AI总结:

提出一种结合LLM的多任务视觉感知框架,通过局部视觉和顺序动作计划实现无全局地图的自主导航,有效减少漂移并提升导航安全性与效率。

AI中文摘要:

服务机器人的长期导航面临关键挑战,如里程计漂移和传感器误差的累积,这会逐渐降低2D地图的质量,并使传统路径规划算法(如A*、RRT*、DiPPer、ViT-A*)随时间推移而失效。为解决此问题,我们提出一个用户友好的交互式框架,消除对全局一致地图的依赖。我们的方法将视觉感知与大型语言模型(LLM)相结合,通过文本或语音解释用户命令。系统不依赖易漂移的全局地图,而是基于局部视觉线索和以自我为中心的几何指令生成顺序动作计划。这些动作计划被顺序执行,使机器人能够在已知和未知环境中安全导航。通过相对于即时目标重置定位,我们的框架以最小累积漂移策略有效工作,确保准确、高效且无碰撞的导航,无需传统地图维护的开销。在真实世界和模拟数据上的实验表明,与其他方法相比有显著改进。我们的源代码在此https URL公开可访问。

英文摘要:

Long-term navigation for service robots faces crit- ical challenges like the accumulation of odometry drift and sensor error, which progressively degrade 2D maps and renders traditional path planning algorithms (e.g., A*, RRT*, DiPPer, ViT-A*) ineffective over time. To address this, we propose a user-friendly, interactive framework that eliminates the reliance on globally consistent maps. Our approach integrates visual perception with Large Language Models (LLM) to interpret user commands via text or voice. Instead of relying on a drift- prone global map, the system generates a sequential action plan based on local visual cues and egocentric geometric instructions. These action plans are executed sequentially, allowing the robot to navigate known and unknown environments safely. By reset- ting localization relative to immediate targets, our framework effectively works with a minimum accumulation drift strategy, ensuring accurate, efficient, and collision-free navigation without the maintenance overhead of traditional mapping. Experiments on real-world and simulated data have shown significant improve- ments over other methods. Our source code is publicly accessible at https://github.com/PraveenSingh24/VL-Navigation.

补充信息

↑