发表机构
Stanford University(斯坦福大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对长时程机器人搜索中自然语言与异质证据的解析难题,提出WORLDS框架,通过持久化图与并行推理器实现结构化世界状态推理,在CityNav上取得51.8%的导航成功率,刷新最佳纪录。
AI 中文摘要
长时程机器人搜索必须将自然语言与异质、不完整且常常模糊的证据进行解析:文本信息、先验地图以及随时间到达的观测。核心挑战在于将这些信息流进行情境化处理,并在选择目标之前决定在哪里收集证据。我们提出了WORLDS:世界状态观测与推理用于语言引导的发现与搜索,这是一个将推理扎根于由地理空间先验初始化并由感知更新的持久化图中的框架。并行推理器维护相互竞争的候选解释,并请求证据以区分它们。我们使用多模态检查器收集和处理所请求的观测,之后由裁判选择有根据的目标或请求另一轮处理。WORLDS在所有5,311个CityNav测试场景中实现了51.8%的导航成功率,这是报告的最高成功率,在仅使用OSM的高分辨率正射影像协议下,比之前发布的最佳结果高出15.7个百分点。在1,000个共享场景中,它实现了50.0%的成功率,而使用相同模型、先验、感知栈和移动预算的最强适配基线为27.9%。检查器的基于观测的验证贡献了其中5.9个百分点的成功,并且在降低推理努力设置下,WORLDS仍比适配的GeoNav基线高出18.8个百分点,同时生成更少的令牌。我们还在一架四旋翼飞行器上演示了WORLDS,该飞行器飞行生成的传感航点,并基于其机载图像确定了三个语言目标,包括一个地图上不存在的车辆。
英文摘要
Long-horizon robotic search must resolve natural language against heterogeneous, incomplete, and often ambiguous evidence: textual information, prior maps, and observations arriving over time. The core challenge is to contextualize these streams and decide where to gather evidence before selecting a target. We present WORLDS: World-state Observation and Reasoning for Language-guided Discovery and Search, a framework that grounds reasoning in a persistent graph initialized from geospatial priors and updated by perception. Parallel Reasoners maintain competing candidate interpretations and request evidence to distinguish between them. We collect and process the requested observations with a multimodal Examiner, after which a Judge selects a grounded target or requests another pass. WORLDS achieves 51.8% navigation success across all 5,311 CityNav test episodes, the highest reported success rate, exceeding the previous published best by 15.7 percentage points under an OSM-only, high-resolution orthographic protocol. On 1,000 shared episodes, it achieves 50.0% versus 27.9% for the strongest adapted baseline using the same model, prior, sensing stack, and movement budget. Observation-based verification by the Examiner contributes 5.9 points of this success, and at a reduced reasoning-effort setting WORLDS still exceeds the adapted GeoNav baseline by 18.8 points while generating fewer tokens. We also demonstrate WORLDS on a quadrotor, which flies the generated sensing waypoints and grounds three language targets, including a vehicle absent from the map, from its onboard imagery.
Comments9 pages, 5 figures, Submitted to the IEEE International Conference on Robotics and Automation (ICRA) 2027