arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

GroupForward:通过实例分组前馈高斯溅射构建可参考的3D场景

GroupForward: Building Referable 3D Scenes via Instance-Grouped Feed-Forward Gaussian Splatting

Qijian Tian, Zimeng Wu, Xuhong Wang, Lizhuang Ma, Xin Tan

arXiv 2608.17535首次发表:更新:

发表机构

School of Computer Science, Shanghai Jiao Tong University; School of Computer Science and Engineering, Beihang University; Shanghai Artificial Intelligence Laboratory; School of Computer Science and Technology, East China Normal University(上海交通大学计算机科学学院; 北京航空航天大学计算机科学与工程学院; 上海人工智能实验室; 华东师范大学计算机科学与技术学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

GroupForward是一种实例分组前馈高斯溅射模型,可从稀疏多视图图像重建3D场景,结合RSRF实现复杂3D参考推理,提升了语义重建与参考推理的性能。

AI 中文摘要

同时重建和理解3D环境对于具身智能体至关重要。为实现这一目标,前馈语义3D高斯溅射(3DGS)可从稀疏多视图观测中高效构建语义场景表示。然而,现有方法缺乏明确的实例判别能力,且主要支持基于类别或短语的语义查询。为此,我们提出GroupForward,这是一种实例分组前馈高斯溅射模型,可从稀疏、无姿态、无校准的多视图图像中重建几何、外观、实例结构和语义。与现有方法将高维语义特征附加到每个高斯的做法不同,GroupForward学习紧凑的实例嵌入,将高斯分组为跨视图一致的3D实例,将前馈语义3DGS从逐高斯语义特征渲染重新表述为实例级语义聚合与传播。基于这些实例组,我们进一步提出用于复杂3D referring分割的参考场景推理框架(RSRF)。RSRF构建实例分组的3D场景图,并为给定的referring表达式检索候选实例。随后,视觉语言模型基于结构化实例证据和多视图观测进行推理,以从候选中识别出被引用的实例。RSRF因此将语言交互从简单的语义查询扩展到复杂的参考场景推理。在语义重建和参考推理任务上的实验验证了我们的实例分组重建与推理框架的有效性。

英文摘要

Simultaneously reconstructing and understanding 3D environments is essential for embodied agents. Toward this goal, feed-forward semantic 3D Gaussian Splatting (3DGS) efficiently constructs semantic scene representations from sparse multi-view observations. However, existing methods lack explicit instance discrimination and mainly support category- or phrase-based semantic queries. To this end, we propose GroupForward, an instance-grouped feed-forward Gaussian splatting model that reconstructs geometry, appearance, instance structure, and semantics from sparse, unposed, and uncalibrated multi-view images. Unlike existing methods that attach high-dimensional semantic features to each Gaussian, GroupForward learns compact instance embeddings that group Gaussians into cross-view consistent 3D instances, reformulating feed-forward semantic 3DGS from per-Gaussian semantic feature rendering to instance-level semantic aggregation and propagation. Building on these instance groups, we further propose a Referential Scene Reasoning Framework (RSRF) for complex 3D referring segmentation. RSRF constructs an instance-grouped 3D scene graph and retrieves candidate instances for a given referring expression. A vision-language model then reasons over structured instance evidence and multi-view observations to identify the referred instance among the candidates. RSRF thereby extends language interaction from simple semantic querying to complex referential scene reasoning. Experiments on semantic reconstruction and referential reasoning demonstrate the effectiveness of our instance-grouped reconstruction and reasoning framework.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑