发表机构
Carnegie Mellon University(卡内基梅隆大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究视觉语言模型可解释零样本分类,指出基于类名的描述符视觉证据不足。提出从目标图像集选属性的方法,该方法提升了准确率,性能优于CoOp,用时短,所选属性还能描述数据分布,是一种有效的分布条件属性选择策略。
AI 中文摘要
一种流行的可解释零样本分类方法是让大语言模型描述每个类名,并将结果描述符用于CLIP。研究表明这些描述符自身视觉证据很少,移除提示中的类名会使ImageNet准确率从59.5%降至15.5%。原因是描述符依赖标签而非图像。因此从目标图像集中选择属性,在CLIP联合嵌入空间中对大属性池与图像评分,按类保留高分属性。无类名属性提示在ImageNet上达到23.8%,在四个转移的ImageNet变体上也有提升,所选属性在单图像时比CoOp方法性能高3分且用时短,还能描述数据分布变化。
英文摘要
A popular route to interpretable zero-shot classification asks a large language model (LLM) to describe each class name and prompts CLIP with the resulting descriptors. We show that these descriptors carry little visual evidence of their own: removing the class name from the prompt collapses ImageNet accuracy from 59.5% to 15.5%. The diagnosis is that the descriptors are conditioned on the label rather than on the images, so they describe the concept in general and mislead exactly when the data shifts; an LLM insists that strawberries are red, but every strawberry in ImageNet-Sketch is a colorless line drawing. We therefore select attributes from the target image collection instead: we score a large attribute pool against the images in CLIP's joint embedding space and keep the top-scoring attributes per class. Selected this way, class-name-free attribute prompts reach 23.8% on ImageNet (against 15.5% for LLM descriptors), the gain holds on four shifted ImageNet variants, and reselecting from the LLM's own pool isolates the selection mechanism as the cause. With one image per class, the selected attributes outperform the prompt-tuning method CoOp by 3 points while fitting in under a minute instead of 14 hours, with no learned soft prompt to obscure the decision. Because the attribute set is chosen by the data, it doubles as a readable summary of a dataset, which we use to describe distribution shift in words. Our code and results are available on our project page: https://ggare-cmu.github.io/AttributeSelect/
CommentsAccepted at the PFATCV Workshop, ECCV 2026. Project page: https://ggare-cmu.github.io/AttributeSelect/