arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

用于细粒度分类的CLIP引导无标签判别区域评分

CLIP-Guided Label-Free Discriminative Region Scoring for Fine-Grained Classification

Yujie Zhu

arXiv 2607.13437首次发表:更新:

发表机构

State University of New York at Buffalo(纽约州立大学布法罗分校)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对细粒度分类中无真实标签时不同区域和评分策略影响不明的问题,提出CLIP引导无标签区域评分框架,用多种策略评估并引入伪标签变体,实验表明软负边际评分性能最佳,随机裁剪在伪标签有噪声时表现更优,为细粒度分类提供新见解。

AI 中文摘要

近期如CLIP和SAM这样的视觉模型实现了无训练分割和细粒度分类的语义编码。常用方法是比较分割图像区域表示与相应标签的文本提示嵌入。但不同局部区域和基于CLIP的评分策略如何影响判别证据选择仍不明,尤其是无真实标签时。本文提出统一的CLIP引导无标签区域评分框架,用SAM生成的掩码和随机裁剪评估多种评分策略,引入两种基于全局和局部嵌入的无标签伪标签变体。在五个数据集上实验比较不同区域生成方法和评分策略。结果表明软负边际评分性能最强,伪标签评分接近真实标签性能。随机裁剪在伪标签有噪声时表现更优,SAM掩码通过聚合所有区域嵌入受益,随机裁剪在较小top-k子集时表现更好。这些发现为细粒度分类提供新见解。

英文摘要

Recent vision models such as CLIP and SAM enable training-free segmentation and semantic encoding for fine-grained classification. A common approach is to compare the representations of segmented image regions with the text prompt embeddings of the corresponding labels. However, it remains unclear how different local regions and CLIP-based scoring strategies affect the selection of discriminative evidence, especially when ground-truth labels are unavailable. In this paper, we propose a unified CLIP-guided label-free region scoring framework for fine-grained classification. The framework evaluates cosine similarity-based, margin-based, and entropy-based scoring strategies using both SAM-generated masks and random crops, and introduces two label-free pseudo-label variants based on global image embeddings and local region embeddings. We conduct experiments on five fine-grained classification datasets to systematically compare different region generation methods and scoring strategies. The results show that Soft Negative Margin scoring achieves the strongest performance, and pseudo-label scoring closely approximates true-label performance. Although SAM produces semantically meaningful masks, random-crop-based pseudo-label scoring consistently outperforms SAM-based scoring across all datasets, suggesting that random crops preserve surrounding information and provide more stable semantic context when pseudo-labels are noisy. In addition, SAM masks benefit from aggregating embeddings from all regions, whereas random crops tend to perform better with a smaller top-k subset. These findings provide new insights for fine-grained classification.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑