arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

3D高斯溅射的后训练语义提升:分离检测器、提升与表示误差

Post-Training Semantic Lifting for 3D Gaussian Splatting: Separating Detector, Lifting and Representation Error

Iván Verdugo Guerra, Ezequiel López Rubio, Jorge García González

arXiv 2610.08756首次发表:更新:

发表机构

University of Málaga(马拉加大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

提出一种后训练语义提升方法,通过多视角证据加权累积和双阈值过滤,分离检测器、提升与表示误差,在ScanNet++上显著提升mIoU。

AI 中文摘要

3D高斯溅射模型的同一个高斯从多个视角被观察,而这些视角对其所属类别并不总是一致。该高斯在其中一些视角中可能被遮挡,并且检测器的置信度在不同视角间也不相同。另一方面,地面真值以带注释的网格形式给出,因为两次训练运行不会产生相同的高斯。在这项工作中,我们提出了一种后训练提升方法,该方法一次处理一个目标类别,并结合来自所有视角的信息。目标和非目标证据同时累积,并根据每个高斯在每个视角中的可见性进行加权。之后,使用两个阈值对高斯进行过滤:主阈值β选择高置信度种子,较低的阈值γβ添加其周围的连通分量。在评估中,标签从高斯转移到既可见又已注释的网格顶点。通过这种设计,我们可以分离三类误差来源:2D检测器、提升以及表示之间的转移。阈值和转移算子在七个Replica验证场景上选择,该方法在十个留出的ScanNet++场景上评估,每个场景和类别使用相同的值。验证场景上的平均mIoU在数据集注释掩码下为0.93,在YOLO掩码下为0.65;在ScanNet++测试场景上分别为0.80和0.54。与之前版本的方法对每个视角的证据进行阈值化相比,该分数将测试mIoU提高了0.24,并使得在两个数据集的所有类别和场景中可以使用单一阈值。最后,误差分析表明大部分剩余误差来自检测器。

英文摘要

The same Gaussian of a 3D Gaussian Splatting model is seen from many views, and these views do not always agree on the class it belongs to. The Gaussian may be occluded in some of them, and the confidence of the detector is not the same from one view to another. The ground truth, on the other hand, is given as an annotated mesh, because two training runs do not produce the same Gaussians. In this work, we propose a post-training lifting method that works with one target class at a time and combines the information coming from all the views. Target and non-target evidence are accumulated simultaneously, weighted by the visibility of each Gaussian in each view. After that, the Gaussians are filtered with two thresholds: a main threshold $β$ selects the high-confidence seeds, and a lower one $γβ$ adds the connected components around them. For the evaluation, the labels are transferred from the Gaussians to the mesh vertices that are both visible and annotated. With this design, we can separate three sources of error: the 2D detector, the lifting and the transfer between representations. The thresholds and the transfer operator are chosen on seven Replica validation scenes, and the method is evaluated on ten held-out ScanNet++ scenes with the same values for every scene and class. The mean mIoU on the validation scenes was 0.93 with masks from the dataset annotations and 0.65 with YOLO masks, and on the ScanNet++ test scenes it was 0.80 and 0.54. Compared with thresholding the evidence per view, as a previous version of the method did, the fraction improves the test mIoU by 0.24 and makes it possible to use a single threshold for all the classes and scenes of both datasets. Finally, the error analysis shows that most of the remaining error comes from the detector.

Comments18 pages, 11 figures, 9 tables. Code: https://github.com/ivanver02/semantic-lifting-3dgs

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑