arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

Seed2GS:通过单个参考视图定位从3D高斯溅射(3DGS)场景中进行无相机、无训练的目标提取

Seed2GS: Camera-Free, Training-Free Object Extraction from 3D Gaussian Scenes via a Single Reference-View Grounding

Zongjian Ding, Yudong Gao, Jiale Liu, Xinglin Yu, Junxing Ren, Dong Wei, Yajing Chen, Shan Huang, Mingjun Cheng, Min Li

arXiv 2608.11928首次发表:更新:

AI 中文总结

Seed2GS是一种无相机、无训练的目标提取方法,通过分离目标身份与3D覆盖范围,在LERF-MASK和3D-OVS数据集上实现了高精度,且计算延迟低。

AI 中文摘要

从预先构建的3D高斯溅射(3DGS)场景中提取目标对象可实现交互式3D编辑。现有方法要么每个场景需训练数十分钟,要么牺牲精度,要么需要预先构建资产可能不包含的原始重建相机。我们提出Seed2GS,它在不使用原始重建相机或场景特定表示训练的情况下,实现了已报道的最高LERF-MASK精度。其核心见解是将目标身份与3D覆盖范围分离。QD-SAM3从多个开放词汇候选中选择一个可靠的参考掩码,一次性固定身份。Seed lift(种子提升)和可见性自适应虚拟轨道随后从新视角暴露对象,同时跟踪传播种子而无需重复检测。由于场景保持冻结,这些掩码仅监督每个高斯的一个临时前景对数。在LERF-MASK上,Seed2GS达到92.1%的平均交并比(mIoU),测得的纯计算延迟为9.3秒,比最强的场景训练基线高3.7个点,比最接近的无相机基线高7.6个点。每个场景使用一个固定的测试参考,完整流程保留91.1%的mIoU;用真实掩码替换其预测种子仅将mIoU提高0.72个点。在3D-OVS上,Seed2GS达到95.7%的mIoU。

英文摘要

Extracting a target object from a pre-built 3D Gaussian Splatting (3DGS) scene enables interactive 3D editing. Existing methods either train for tens of minutes per scene, sacrifice accuracy, or require original reconstruction cameras that pre-built assets may not include. We present Seed2GS, which achieves the highest reported LERF-MASK accuracy without original reconstruction cameras or scene-specific representation training. Its key insight is to separate target identity from 3D coverage. QD-SAM3 selects one reliable reference mask from several open-vocabulary candidates, fixing identity once. Seed lift and visibility-adaptive virtual orbits then expose the object from new viewpoints, while tracking propagates the seed without repeated detection. Because the scene remains frozen, these masks supervise only one temporary foreground logit per Gaussian. On LERF-MASK, Seed2GS reaches 92.1% mean intersection over union (mIoU) with a measured compute-only latency of 9.3 seconds, 3.7 points above the strongest scene-trained baseline and 7.6 points above the closest camera-free baseline. With one fixed test reference per scene, the complete pipeline retains 91.1% mIoU; replacing its predicted seed with a ground-truth mask improves mIoU by only 0.72 points. On 3D-OVS, Seed2GS reaches 95.7% mIoU.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑