arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

DisParQ:用于可解释视觉基础模型的自监督部件概念

DisParQ: Self-Supervised Part Concepts for Interpretable Vision Foundation Models

Adam Pardyl, Siddhartha Gairola, Sukrut Rao, Adam Wróbel, Bartosz Zieliński, Bernt Schiele, Dawid Rymarczyk

arXiv 2610.09802首次发表:更新:

发表机构

Jagiellonian University; Max Planck Institute for Informatics; Ardigen SA(雅盖隆大学; 马克斯·普朗克信息学研究所; Ardigen公司)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

DisParQ通过自监督学习从冻结的视觉骨干中提取离散部件概念,无需标签或语言,在ImageNet线性探测达到83.2%的top-1准确率,并支持跨类别部件检索。

AI 中文摘要

基于概念的视觉模型通过人类可检查概念的中间层来表示图像,因此模型所依赖的内容可以追溯到这些概念。然而,这些模型通常局限于固定类别或依赖语言来定义其概念。我们提出了DisParQ(具有量化属性的离散部件),一种从强大的冻结视觉仅自监督骨干网络中学习空间定位、离散概念表示的方法。它不需要类别标签,也不需要语言监督。每个图像块被分配到一个可学习的原型字典中的恰好一个概念,并且每张图像只能激活概念的稀疏子集。为了捕捉每个概念在不同图像中的变化(例如,“轮子”的类型),我们在概念旁边学习连续残差,然后将其量化为离散属性。空间解码器仅从概念和属性重建骨干网络的表示,因此成功的重建意味着离散表示保留了骨干网络的信息。我们在七个数据集上评估了DisParQ,从通用识别(ImageNet、PartImageNet、Places)到细粒度基准(CUB、Cars、Dogs、Flowers)。我们表明,DisParQ在ImageNet线性探测上与其冻结的DINOv2教师模型紧密匹配(top-1准确率83.2%),实现了比语言对齐模型更高的概念一致性,在细粒度识别上保持竞争力,并支持跨类别的基于部件的检索。

英文摘要

Concept-based vision models represent images through an intermediate layer of human-inspectable concepts, so what a model relies on can be traced to those concepts. However, those models are often limited to fixed categories or depend on language to define their concepts. We introduce DisParQ (Discrete Parts with Quantized attributes), a method that learns spatially grounded, discrete concept representations from a powerful frozen vision-only self-supervised backbone. It requires no class labels and no language supervision. Each image patch is assigned to exactly one concept from a learnable prototype dictionary, and only a sparse subset of concepts may activate per image. To capture how each concept varies across images (e.g., the type of a "wheel"), we learn continuous residuals alongside the concepts and then quantize them into discrete attributes. A spatial decoder reconstructs the backbone's representation from the concepts and attributes alone, so successful reconstruction means that the discrete representation preserves the backbone's information. We evaluate DisParQ across seven datasets, from general recognition (ImageNet, PartImageNet, Places) to fine-grained benchmarks (CUB, Cars, Dogs, Flowers). We show that DisParQ closely matches its frozen DINOv2 teacher on ImageNet linear probing (83.2% top-1), achieves higher concept consistency than language-aligned models, remains competitive on fine-grained recognition, and enables cross-category part-based retrieval.

CommentsUnder review

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑