arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

RECAST:将视觉语言语义重铸为机器人导航的可执行代价地图

RECAST: Recasting Vision-Language Semantics into an Actionable Cost Map for Robot Navigation

Incheol Cho, Jintae Park, Jinkyu Kim, Jungbeom Lee, Jaegul Choo, Seokha Moon

arXiv 2609.32595首次发表:更新:

AI 中文总结

RECAST提出一种机器人导航框架,结合视觉语言模型推理与视觉基础模型空间锚定,构建可执行代价地图,显著提升仿真和真实场景的导航成功率并降低碰撞率。

AI 中文摘要

在多样环境中进行安全且稳健的机器人导航,需要对复杂场景有高层次的理解,并能够将其转化为稳定的运动。近期的工作通过大规模训练的学习模型以及基于视觉语言模型(VLM)的方法来解决这一问题。然而,学习模型在其训练分布之外会失效,而基于VLM的方法虽然带来了这种理解,但很少将其锚定在场景中或使动作与之对齐。为了解决这些局限,我们提出了RECAST,一个机器人导航框架,它将VLM的推理能力与视觉基础模型的空间锚定能力相结合,以构建可执行的代价地图。给定机器人的前视图和用户的指令,我们首先利用VLM对场景进行分解,判断哪些表面是可通行的,哪些物体构成风险,哪个朝向是优选的,以及哪些间隙可以通过。随后,视觉基础模型将这些表面和物体在图像中锚定,所有四个判断被空间重铸为一个紧凑的代价地图。该地图既条件化轨迹解码器,又对其提议进行评分以选择要执行的轨迹。由于VLM的回答滞后于实时场景,两个步骤都利用来自两个时间点的代价地图:VLM所判断的枢轴帧(承载所有四个判断)以及当前帧(其地形和碰撞代价从当前图像重建)。RECAST在仿真中将成功率比最强先前方法提高了13.3个百分点,在真实四足机器人上提高了31.4个百分点,并将碰撞率相对降低了9.6和14.3个百分点,达到所有方法中最低的碰撞率。项目页面可从此https URL访问。

英文摘要

Safe and robust robot navigation across diverse environments requires a high-level understanding of complex scenes and the ability to carry it into stable motion. Recent works tackle this with learning-based models trained at scale and with approaches built on vision-language models (VLMs). However, learning-based models break down outside their training distribution, while VLM-based approaches bring that understanding but rarely ground it in the scene or align the action with it. To address these limitations, we present RECAST, a robot navigation framework that combines the reasoning of a VLM with the spatial grounding of vision foundation models to build an Actionable Cost map. Given the robot's front view and the user's instruction, we first decompose the scene with the VLM, judging which surfaces are traversable, which objects pose a risk, which heading to prefer, and which gaps are passable. Vision foundation models then ground these surfaces and objects in the image, and all four judgments are spatially recast into one compact cost map. This map both conditions the trajectory decoders and scores their proposals to select the one to execute. As the VLM's answers trail the live scene, both steps draw on cost maps from two points in time: the pivot frame the VLM judged, which carries all four judgments, and the current frame, whose terrain and collision costs are rebuilt from the current image. RECAST improves success over the strongest prior method by 13.3 points in simulation and 31.4 points on a real quadruped, and reduces the collision rate relative to it by 9.6 and 14.3 points, reaching the lowest collision rate among all methods. The project page is available at https://recast-nav.github.io/

Comments8 pages, 4 figures. Project page: https://recast-nav.github.io/

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑