AI 中文总结
研究针对机器人社交导航问题,提出G2-Nav框架,将VLM语义推理转化为视觉语言代价地图,经开放集感知评估区域和代理,结合语义验证与高频安全检查,实现非结构化环境下安全、高效且符合社会规范的自主导航。
AI 中文摘要
社交导航要求机器人在复杂的现实世界环境中进行推理和响应。虽然最近的工作试图使用大型视觉语言模型(VLM)将人类级别的智能纳入机器人规划,但端到端框架往往会创建一个不可预测的黑箱,并且现有的指令跟随方法并非为完全自主而设计。为了弥合这一差距,我们提出了G2-Nav,这是一个新颖的框架,它将抽象的社会推理与安全的现实世界部署相结合。G2-Nav不是要求VLM做出直接规划决策,而是将其语义推理转化为具有可靠性和可解释性的视觉语言代价地图。VLM从开放集感知中评估可通行区域和社会代理,将社会背景映射到代价地图中。为了提高现实世界的鲁棒性,VLM对上游跟踪进行语义验证,并且我们引入了高频安全检查以在轨迹生成之前防范系统延迟。我们通过现实世界实验证明,G2-Nav在非结构化环境中提供安全、高效且符合社会规范的自主导航。代码将公开提供。
英文摘要
Social navigation requires the robot to reason and respond in complex real-world environments. While recent works attempt to incorporate human-level intelligence into robot planning using large Vision-Language Models (VLMs), end-to-end frameworks often create an unpredictable black-box, and existing instruction-following methods are not designed for full autonomy. To bridge this gap, we present G2-Nav, a novel framework that grounds abstract social reasoning and guards safe real-world deployment. Instead of asking the VLM for direct planning decisions, G2-Nav translates its semantic reasoning into a vision-language costmap with reliability and interpretability. The VLM evaluates traversable regions and social agents from open-set perception, mapping social context into the costmap. To improve real-world robustness, the VLM performs semantic verification on upstream tracking, and we introduce a high-frequency safety check to guard against system latency prior to trajectory generation. We demonstrate through real-world experiments that G2-Nav delivers safe, efficient, and socially compliant autonomous navigation in unstructured environments. Code is available at https://github.com/centiLinda/G2-Nav.
CommentsCoRL 2026