arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

当外观失效时,几何发挥作用:一种可补充视觉基础模型的无CAD 3D形状先验

What Do Scan-Derived Class Prototypes Add? Disentangling Supervision, Prototype Content and Query Protocol in Recognition over Frozen Foundation Features

Chenxi Tao, Hong-In Won, Seung-Kyum Choi

arXiv 2609.04381首次发表:更新:

发表机构

Georgia Institute of Technology(佐治亚理工学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究提出一种无CAD的3D形状先验,将3D高斯溅射重建的形状原型与DINOv2特征融合,可补充视觉基础模型,提升几何相似低纹理对象的识别性能,尤其在部分遮挡下效果显著。

AI 中文摘要

在制造和服务机器人领域,识别未通过带标签训练集接入的特定对象是一个反复出现的需求,但传统的可渲染先验——计算机辅助设计(CAD)模型——往往无法获取。二维采集无法提供形状先验,而冻结的基础特征在几何相似、低纹理的工业部件上表现不佳。我们探究以对象为中心的短扫描能为识别提供超出采集图像本身的什么价值:每个对象用3D高斯溅射(3DGS)重建,总结为每类形状原型,并与冻结的DINOv2图像特征融合。首先,这种扫描无需CAD即可恢复CAD的识别价值:在T-LESS数据集上,来自RGB-D深度、3DGS和CAD的几何给出的识别结果具有可比性(在HOPE数据集上持平,在T-LESS数据集上差值在1.6个点以内);3DGS只是获取点云的便捷途径。其次,收益取决于形状的可识别性:在形状独特的家庭对象(HOPE)上,仅几何达到0.920,仅图像为0.832,固定权重融合(0.872)处于两者之间;在形状易混淆的无纹理工业部件(T-LESS)上,收益适中但稳定(融合后从0.560升至0.591,高于两种单一信号)。第三,该先验是互补的,而非均匀叠加:它挽救的图像失败远多于破坏的成功,且其益处随部分遮挡而增加。最后,价值在于几何而非渲染像素:3DGS渲染对图像侧无帮助,且冻结特征识别几乎与光照无关(差值在2.5个点以内)。本研究范围限于识别,不涉及BOP姿态基准。

英文摘要

A scan supplies labeled images and a geometric reference. We separate their contributions in a recognizer whose scan-derived prototype matrix acts as a supervised head's fixed output layer. On T-LESS, HOPE and 18 self-collected industrial parts, we test real, random and exactly permuted prototypes, matched geometry-free classifiers, stronger appearance rules and paired background protocols. Across DINOv2-giant and MetaCLIP-H with real-background queries, the largest fused-accuracy advantage of the real prototypes over either control is one percentage point; larger differences favor controls, by up to 2.8 points in arm means. On HOPE with DINOv2-giant the head alone is 2.8 points above exact permutations (95% interval: 0.8-4.7); this advantage does not reach fusion and is not observed on MetaCLIP-H. On DINOv2-giant, matched logistic regression comes within 0.5 points of fusion on T-LESS and exceeds it on HOPE and the self-collected parts. Against white cutouts, real HOPE query backgrounds lower image-prototype accuracy by 43 points on DINOv2-giant and 13 on MetaCLIP-H. The audit separates prototype content, label supervision and query protocol.

Comments35 pages, 7 figures, 14 tables. Revised version with a new title; adds prototype controls, matched supervision references, a second backbone, paired query protocols, a third dataset and an external experiment on Hyperspherical Prototype Networks

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑