arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.20586cs.ROcs.CVeess.IV

CoRef-GS:面向多智能体场景理解的协同指代高斯泼溅

CoRef-GS: Cooperative Referring Gaussian Splatting for Multi-Agent Scene Understanding

Zhikun Zhou, Kunyu Peng, Runyi Yang, Junhao Cai, Di Wen, Ruiping Liu, Danda Pani Paudel, Yi Zhou, Luc Van Gool, Kailun Yang

首次发表
浏览论文内容

中文总结 AI 辅助

针对多智能体协同指代场景理解,提出CoRef-GS框架,通过构建开放词汇实例感知高斯地图、跨智能体对齐和视角条件关系图实现查询接地,并引入双四足基准,显著提升指代精度。

中文摘要 AI 辅助

具身机器人的指代场景理解要求从指定视角对以物体和关系为中心的语言查询进行接地。虽然局部语义高斯地图可以在单个智能体的观测范围内支持这种接地,但协同设置要求在地图独立重建并对齐融合后,该能力仍然有效。在此设置中,指代目标或其上下文地标可能来自另一智能体的观测,而空间关系仍须从查询机器人的视角进行解释。我们将此问题形式化为融合地图上的协同指代高斯接地,这需要几何可对齐性、实例级语义可比性以及视角条件的关系推理。现有的语言感知高斯方法主要关注单地图查询,而高斯配准方法优化几何或光度对齐,却不保留面向语言接地的语义兼容性。我们提出CoRef-GS,一种协同指代高斯泼溅框架。CoRef-GS构建局部开放词汇的实例感知高斯地图,然后通过跨智能体对齐模块利用几何和语义一致性对齐部分重叠的地图,并使用视角条件的掩码关系图对接查询。我们进一步引入CoQuad-Ref,一个涵盖真实世界和模拟室内场景的双四足机器人基准。实验表明,在模拟场景中,CoRef-GS将旋转误差从粗初始化后的2.58°降至细化后的0.15°,并将真实世界指代mIoU从ReferSplat的52.6%提升至68.8%。所建立的基准和源代码将在该https URL公开发布。

英文摘要

Referring scene understanding for embodied robots requires grounding object- and relation-centric language queries from a designated viewpoint. While a local semantic Gaussian map can support such grounding within one agent's observations, cooperative settings require this ability to remain effective after independently reconstructed maps are aligned and fused. In this setting, the referred target or its contextual landmark may come from another agent's observations, while spatial relations must still be interpreted from the querying robot's viewpoint. We formulate this problem as cooperative referring Gaussian grounding over fused maps, which requires geometric alignability, instance-level semantic comparability, and view-conditioned relation reasoning. Existing language-aware Gaussian methods mainly focus on single-map querying, whereas Gaussian registration methods optimize geometric or photometric alignment without preserving language-grounding-oriented semantic compatibility. We propose CoRef-GS, a cooperative referring Gaussian splatting framework. CoRef-GS constructs local open-vocabulary instance-aware Gaussian maps, then aligns partially overlapping maps with a cross-agent alignment module by geometric and semantic consistency, and grounds queries using a view-conditioned mask relation graph. We further introduce CoQuad-Ref, a dual-quadruped benchmark spanning both real-world and simulated indoor scenes. Experiments show that, on simulated scenes, CoRef-GS reduces the rotation error from 2.58° after coarse initialization to 0.15° after refinement, and improves real-world referring mIoU over ReferSplat from 52.6% to 68.8%. The established benchmark and source code will be publicly released at https://github.com/ruojiruoli17/CoRef-GS.git.

发表机构

  • Hunan University(湖南大学)
  • National Engineering Research Center of Robot Visual Perception and Control Technology(国家机器人视觉感知与控制技术工程研究中心)
  • Karlsruhe Institute of Technology(卡尔斯鲁厄理工学院)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑