发表机构
Southeast University; The Hong Kong Polytechnic University; Tsinghua University(东南大学; 香港理工大学; 清华大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
NavPatch利用视觉语言模型进行场景理解,对物体类别执行ADD、REMOVE或EXTEND操作以修正代价地图,在50次真实试验中成功率86.0%,较基线提升16个百分点。
AI 中文摘要
移动机器人通常依赖几何地图进行避障和路径规划,但由此产生的障碍物表示并不总能与物体对导航的影响方式相匹配。低矮的电缆可能被遗漏,柔性窗帘可能造成虚假的阻塞,而交通锥可能需要比其观测足迹更大的排除区域。我们提出了NavPatch,一种物体级修正层,通过视觉语言模型进行周期性场景理解,为导航相关物体类别分配ADD、REMOVE或EXTEND操作。开放词汇定位可定位物体实例,LiDAR和RGB-D观测提供3D支撑。观测质量过滤和跨帧维护决定每个修正补丁何时被提交、替换或撤销。在跨越五种布局的50次真实机器人试验中,NavPatch实现了86.0%的总体成功率。对四种配置共200次运行的消融研究表明,与仅基于当前观测的更新相比,NavPatch将成功率从70.0%提升至86.0%,并将错误提交率从68.4%降低至40.7%。
英文摘要
Mobile robots typically rely on geometric maps for obstacle avoidance and path planning, but the resulting obstacle representation does not always match how an object should affect navigation. A low lying cable may be missed, a flexible curtain may create spurious blockage, and a traffic cone may require an exclusion region larger than its observed footprint. We present NavPatch, an object level correction layer that assigns ADD, REMOVE, or EXTEND to navigation relevant object categories through periodic scene understanding with a vision-language model. Open vocabulary grounding localizes object instances, and LiDAR and RGB-D observations provide 3D support. Observation quality filtering and cross frame maintenance determine when each correction patch is committed, replaced, or revoked. In 50 real robot trials across five layouts, NavPatch achieves an overall success rate of 86.0%. An ablation study of four configurations with 200 runs in total shows that NavPatch improves the success rate from 70.0% to 86.0% and reduces the false commit rate from 68.4% to 40.7% compared with updates based only on the current observation.