发表机构
Beihang University; The University of Hong Kong(北京航空航天大学; 香港大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
提出DIFTA-3D,通过深度一致特征管线将DINOv3适配到IIFNet3D实例级融合,在ScanNetV2上以Conservative VAID蒸馏获得76.59/62.16 mAP,验证了受控迁移方案的有效性。
AI 中文摘要
RGB-D 3D实例检测器受益于视觉语义,但IIFNet3D使用的任务特定的Faster R-CNN/ResNet分支将特征提取与单独训练的2D检测器及其图像域标签耦合在一起。用冻结的视觉基础模型替换该分支可以消除这种任务特定的依赖,但可能引入遮挡噪声以及patch特征与几何感知检测特征之间的不匹配。在本工作中,我们通过将DINOv3适配到IIFNet3D的实例级融合流程来研究这种替换。我们方法的核心是一个深度一致的特征管线,该管线将场景点投影到校准的RGB-D帧中,应用度量深度残差检查,将接受的DINOv3特征平均到离线点缓存中,并在提议对齐的RoI网格内聚合缓存特征。几何和双向实例融合路径得以保留,同时将Conservative VAID评估为一种仅应用于正RoI的低强度、支持加权的语义蒸馏方案。我们在ScanNetV2上进行了广泛评估以评估所提出的迁移方案。在ScanNetV2上,我们的DINOv3对照组在IoU阈值为0.25和0.50时分别达到76.15和60.93的mAP分数。Conservative VAID设置在检查点级方案比较中分别达到76.59和62.16的mAP分数,相对于对照组分别获得0.44和1.23个点的数值提升。报告的IIFNet3D结果75.7/63.8仅用作外部参考,因为视觉分支和处理协议不同。因此,我们将这些结果解释为受控迁移方案的证据,而非VAID或深度过滤各自贡献的因果估计。
英文摘要
RGB-D 3D instance detectors benefit from visual semantics, but the task-specific Faster R-CNN/ResNet branch used by IIFNet3D couples feature extraction to a separately trained 2D detector and its image-domain labels. Replacing that branch with a frozen vision foundation model removes this task-specific dependency, but may introduce occlusion noise and a mismatch between patch features and geometry-aware detection features. In this work, we investigate this replacement through an adaptation of DINOv3 to the instance-level fusion pipeline of IIFNet3D. At the core of our approach is a depth-consistent feature pipeline that projects scene points into calibrated RGB-D frames, applies a metric depth-residual check, averages the accepted DINOv3 features into an offline point cache, and aggregates the cached features inside proposal-aligned RoI grids. The geometric and bidirectional instance-fusion paths are preserved, while Conservative VAID is evaluated as a low-strength, support-weighted semantic distillation recipe applied only to positive RoIs. We conduct extensive evaluations on ScanNetV2 to assess the proposed transfer recipes. On ScanNetV2, our DINOv3 control achieves mAP scores of 76.15 and 60.93 at IoU thresholds of 0.25 and 0.50, respectively. The Conservative VAID setting achieves mAP scores of 76.59 and 62.16, corresponding to numerical gains of 0.44 and 1.23 points over the control, respectively, in this checkpoint-level recipe comparison. The reported IIFNet3D result of 75.7/63.8 is used only as an external reference because the visual branch and processing protocol differ. Accordingly, we interpret these results as evidence for a controlled transfer recipe rather than as a causal estimate of the individual contributions of VAID or depth filtering.
Comments9 pages, 6 figures, conference paper