arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

DSSR-3D:三维高斯中视角依赖指代的解耦推理

DSSR-3D: Decoupled Reasoning for View-Dependent Referring in 3D Gaussians

Thanh-Khoi Nguyen, Thien-Phuc Tran, Minh-Triet Tran

arXiv 2610.00040首次发表:更新:

发表机构

University of Science, Ho Chi Minh City; Viet Nam National University, Ho Chi Minh City(胡志明市理科大学; 越南国立大学胡志明市分校)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

提出DSSR-3D,通过解耦语义定位与空间推理的推理时框架,在三维高斯场上实现视角依赖指代分割,无需重训练,并构建ViewRef-GS基准验证其零样本迁移与性能提升。

AI 中文摘要

近期三维高斯泼溅技术的进展,通过将二维基础模型中的语义知识蒸馏到三维表示中,实现了开放词汇和指代分割。然而,现有的指代场将语言特征嵌入到全局视角不变的空间中,从根本上无法解决依赖于相机姿态的以观察者为中心的空间关系(例如“在……左侧”)。我们提出DSSR-3D,一种在连续三维高斯场上进行视角依赖指代分割的推理时框架,将其形式化为两个接口——姿态不变的语义定位和姿态条件下的空间推理——使得满足这些约束的任意函数对都能产生有效的实例化,无需重新训练底层语义场,也不依赖离散的几何代理(如边界框)。我们用温度锐化的softmax定位机制和基于投影的方向评分函数实例化这两个接口,通过轻量级、无需训练的步骤融合,并展示它们无需适应即可零样本迁移到结构不同的语义场。我们进一步提出ViewRef-GS,一个在三维高斯场上隔离视角依赖分割的基准,与增强的Ref-LERF联合评估,为视角依赖的空间接地提供全面的测试平台。实验表明,在基础语义场之外无需额外训练的情况下,相较于现有基于3DGS的指代方法,取得了持续的性能提升。

英文摘要

Recent advances in 3D Gaussian Splatting have enabled open-vocabulary and referring segmentation by distilling semantic knowledge from 2D foundation models into 3D representations. However, existing referring fields embed language features in a globally view-invariant space, making them fundamentally unable to resolve observer-centric spatial relations (e.g., "to the left of") that depend on camera pose. We propose DSSR-3D, an inference-time framework for view-dependent referring segmentation on continuous 3D Gaussian fields, formalized as two interfaces - pose-invariant semantic localization and pose-conditioned spatial reasoning - such that any pair of functions satisfying these constraints yields a valid instantiation, requiring no retraining of the underlying semantic field and no reliance on discrete geometric proxies such as bounding boxes. We instantiate the two interfaces with a temperature-sharpened softmax localization mechanism and a projection-based directional scoring function, fused via a lightweight, training-free step, and show they transfer zero-shot to structurally distinct semantic fields without adaptation. We further propose ViewRef-GS, a benchmark isolating view-dependent segmentation on 3D Gaussian fields, evaluated jointly with an augmented Ref-LERF to provide a comprehensive testbed for viewpoint-dependent spatial grounding. Experiments show consistent gains over existing 3DGS-based referring methods, with no additional training beyond the base semantic field

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑