arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

CALIPER:基于度量锚定的无模型识别视觉相似工业零件

CALIPER: Metric-Grounded Model-Free Recognition of Visually Similar Industrial Parts

Alankrit Gupta, Chenxi Tao, Seung-Kyum Choi

arXiv 2609.17820首次发表:更新:

发表机构

Georgia Institute of Technology(佐治亚理工学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

CALIPER提出一种无模型RGB-D框架,结合支持匹配与度量尺寸证据,实现视觉相似工业零件的细粒度识别,在18类零件上达到88.2%闭集准确率,并支持免训练注册新类别。

AI 中文摘要

当类别主要差异在于物理尺寸时,对视觉相似的工业零件进行细粒度识别具有挑战性。将检测到的物体裁剪区域归一化到固定输入尺寸会抑制绝对尺度信息,而在不断演进的工业库存中,CAD模型和大型类别特定数据集可能不可用。我们提出CALIPER,一种无模型的RGB-D框架,将基于支持集的表观匹配与度量尺寸证据相结合。每个训练类别通过单个转盘RGB-D视频和一到两张标注的真实图像完成注册;3D重建提供新视角的表观支持,而对齐的深度图生成类别特定的度量尺寸轮廓。在推理时,粗粒度YOLOv8n-seg模型定位零件,冻结的DINOv2骨干网络配合情景训练(episodically trained)的嵌入头执行细粒度支持匹配。基于边际条件的度量融合仅在表观模糊的决策中激活概率性尺寸证据。新类别通过小型RGB-D支持集注册,无需更新网络参数。我们在18个视觉相似的工业零件上评估CALIPER:其中16个类别用于训练,而两个螺钉保留用于免训练注册。CALIPER在闭集上达到88.2%的准确率,定位召回率为99.8%,在对两个未见螺钉进行10次样本注册后,总体准确率达到85.7%。度量融合将未见类别的准确率提升了最多37.4个百分点,且对原有库存没有统计显著的性能下降。机器人臂部署无需针对部署环境的重新训练即可识别17/18个零件。

英文摘要

Fine-grained recognition of visually similar industrial parts is challenging when classes differ primarily in physical dimensions. Normalizing detected object crops to a fixed input size suppresses absolute scale, while CAD models and large class-specific datasets may be unavailable in evolving industrial inventories. We present CALIPER, a model-free RGB-D framework that couples support-based appearance matching with metric size evidence. Each training class is onboarded from a single turntable RGB-D video and one to two labeled real images; 3D reconstruction provides novel-view appearance support, while aligned depth yields a class-specific metric size profile. At inference, a coarse YOLOv8n-seg model localizes parts, and a frozen DINOv2 backbone with an episodically trained embedding head performs fine-grained support matching. Margin-conditioned metric fusion activates probabilistic size evidence only for appearance-ambiguous decisions. New classes are enrolled from a small RGB-D support set without updating network parameters. We evaluate CALIPER on 18 visually similar industrial parts: 16 classes are used for training, while two screws are reserved for training-free enrollment. CALIPER achieves 88.2% closed-set accuracy with 99.8% localization recall and 85.7% overall accuracy after 10-shot enrollment of the two unseen screws. Metric fusion improves unseen-class accuracy by up to 37.4 percentage points without statistically significant degradation of the original inventory. Robot-arm deployment identifies 17/18 parts without deployment-specific retraining.

Comments5 figures, 5 tables

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑