arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.07144cs.CV

InstanceSplat:面向场景理解的实例感知前馈三维高斯溅射

InstanceSplat: Instance-Aware Feed-Forward 3D Gaussian Splatting for Scene Understanding

Minchao Jiang, Xiaoxuan Ma, Shunyu Jia, Haoru Wang, Zhang Liang, Wentao Zhu

首次发表
浏览论文内容

中文总结 AI 辅助

InstanceSplat是统一前馈3DGS框架,通过单次前向传播构建联合编码多信息的实例感知高斯表示,以实例为中心的学习策略实现重建与场景理解互益,在多任务上达到SOTA性能与强泛化性。

中文摘要 AI 辅助

前馈三维高斯溅射(3DGS)可实现高效且泛化性强的三维重建,但当前用于场景理解的前馈3DGS方法大多以类别为导向。相比之下,实例感知的3DGS方法通常依赖逐场景优化,且常将重建与实例、语义学习解耦,限制了三者间的相互作用。我们提出InstanceSplat,这是一个统一的前馈3DGS框架,用于从无姿态多视图图像中实现泛化性三维重建和实例感知场景理解。在单次前向传播中,InstanceSplat构建了一个实例感知的高斯表示,该表示联合编码外观、几何、实例身份和语言对齐的语义。共享的3D高斯为跨视图的实例身份提供基础,生成可渲染且跨视图一致的实例特征。为使重建和场景理解能相互受益,我们进一步设计了以实例为中心的学习策略,通过共享实例结构将重建、实例学习和语义学习连接起来。具体而言,实例线索引导重建,语言对齐的语义增强易混淆同类实例的区分度,实例区域将语义证据聚合为连贯的对象级预测。在不同输入视图设置下及未见过的数据集上,针对新视图合成、实例分割和开放词汇语义理解的实验表明,该方法达到了最先进的性能、实际效率和强泛化能力。

英文摘要

Feed-forward 3D Gaussian Splatting (3DGS) enables efficient and generalizable 3D reconstruction, but current feed-forward 3DGS methods for scene understanding remain largely category-oriented. In contrast, instance-aware 3DGS methods typically rely on per-scene optimization and often decouple reconstruction from instance and semantic learning, limiting reciprocal interactions among them. We present InstanceSplat, a unified feed-forward 3DGS framework for generalizable 3D reconstruction and instance-aware scene understanding from pose-free multi-view images. In a single forward pass, InstanceSplat constructs an instance-aware Gaussian representation that jointly encodes appearance, geometry, instance identity, and language-aligned semantics. Shared 3D Gaussians ground instance identities across views, producing renderable and cross-view-consistent instance features. To allow reconstruction and scene understanding to benefit from each other, we further design an instance-centric learning strategy that connects reconstruction, instance learning, and semantic learning through shared instance structure. Specifically, instance cues guide reconstruction, language-aligned semantics strengthen the discrimination of confusing same-category instances, and instance regions aggregate semantic evidence into coherent object-level predictions. Experiments on novel-view synthesis, instance segmentation, and open-vocabulary semantic understanding under varying input-view settings and on an unseen dataset demonstrate state-of-the-art performance, practical efficiency, and strong generalization.

补充信息

↑