发表机构
University of Waterloo; Sun Yat-sen University; University of Calgary(滑铁卢大学; 中山大学; 卡尔加里大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
CDSeg是基于高斯溅射的跨域分割接口,无需特定任务3D分割训练,可复用2D掩码实现多视图标签迁移,在多个数据集上取得高mIoU,能快速处理大规模场景。
AI 中文摘要
现代图像模型能为每个视图中应分割的内容提供强提示,但其生成的掩码本身无法确定这些标签在3D空间中的保留位置。本文提出跨域分割方法CDSeg(基于高斯溅射技术),这是一种无需特定任务3D分割训练的标签迁移接口,采用高斯基元作为可渲染标签载体。外部掩码源提供标签,渲染器生成的可见性决定哪些3D基元接收标签;该载体可通过将每个输入点补全为一个高斯并保留其索引来实例化,或复用优化后高斯场景的原生基元。CDSeg在渲染过程中记录像素与基元的关联,通过投票和局部滤波融合多视图掩码,生成的标签可返回至原始点、保留在原生高斯场景中或渲染到其他视图。CDSeg支持可提示、自动实例、语义及LiDAR设置,能在数秒内处理含数百万基元的场景;在DesktopObjects-360数据集上的mIoU达92.35%,NeRDS-360数据集上达95.89%,使用提供的2D语义标注在完整ScanNet-v2验证集上达65.77%。CDSeg为跨点云、高斯场景及图像视图复用2D掩码提供统一接口,无需特定任务的3D分割网络。
英文摘要
Modern image models provide strong cues about \emph{what} should be segmented in each view, but their masks do not by themselves determine \emph{where} those labels should persist in 3D. We present Cross-Domain Segmentation via Gaussian Splatting (CDSeg), a label-transfer interface that requires no task-specific 3D segmentation training and uses Gaussian primitives as a renderable label carrier. An external mask source supplies the labels, while renderer-derived visibility determines which 3D primitives receive them. The carrier is instantiated either by completing each input point into one Gaussian, preserving its index, or by reusing the native primitives of an optimized Gaussian scene. CDSeg records pixel--primitive associations during rendering and fuses multi-view masks through voting and a local filter. The resulting labels can be returned to the original points, retained on the native Gaussian scene, or rendered into other views. CDSeg covers promptable, automatic instance, semantic, and LiDAR settings and processes scenes with millions of primitives in seconds. It obtains 92.35\% mIoU on DesktopObjects-360, 95.89\% on NeRDS-360, and 65.77\% on the full ScanNet-v2 validation split using the provided 2D semantic annotations. CDSeg thereby provides one interface for reusing 2D masks across point clouds, Gaussian scenes, and image views without a task-specific 3D segmentation network.
Comments15 pages, 8 figures