发表机构
University of Melbourne; Southeast University; RIKEN Center for Advanced Intelligence Project; The University of Tokyo(墨尔本大学; 东南大学; 理化学研究所先进智能项目中心; 东京大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究针对预训练视觉语言模型零样本图像分类中提示与类独立性假设不成立的问题,提出类感知零样本提示重加权方法CARPRT,通过无训练方式量化提示与类的相关性并调整权重,实验表明该方法优于现有方法。
AI 中文摘要
预训练的视觉语言模型(VLMs)通过计算图像与文本描述之间的相似性分数来实现零样本图像分类,文本描述通常是将类标签插入提示中形成的。由于给定图像-类对的分数对提示选择敏感,现有研究使用加权向量对多个提示进行集成以聚合不同提示的分数。然而,当前策略中分配给每个提示的加权向量在所有类之间共享,隐含地假设提示与类条件独立,这在实践中往往不成立。为解决此问题,我们提出了类感知零样本提示重加权(CARPRT)。这种评分方案通过以无训练的方式捕获不同提示的类特定相关性来调整每个类标签的加权向量。对于每个类标签和每个可用提示,我们通过对在给定提示下预测到该类的图像上的图像-文本相关性分数求平均来量化它们的类特定相关性。然后对这些估计进行归一化以得出类特定权重。在标准图像分类基准上的评估表明,CARPRT优于现有的类独立重加权方法,证实了对提示-类依赖性进行建模对于有效的零样本预测以及更广泛的基于VLM的依赖提示集成的应用设置至关重要。我们的代码可在这个https URL上获取。
英文摘要
Pre-trained vision-language models (VLMs) enable zero-shot image classification by computing the similarity score between an image and textual descriptions, typically formed by inserting a class label (e.g., "cat") into a prompt (e.g., "a photo of a"). Since the score for a given image-class pair is sensitive to the choice of prompt, existing studies ensemble multiple prompts using a weighting vector to aggregate scores across different prompts. Yet, in current strategies, the weighting vector assigned to each prompt is shared across all classes, implicitly assuming that prompts are conditionally independent of classes, which often does not hold in practice, as a prompt like "an aerial view of" might be apt for "airport" but ill-suited for "apple". To address this, we propose class-aware zero-shot prompt reweighting (CARPRT). This scoring scheme adjusts the weighting vector for each class label by capturing the class-specific relevance of different prompts in a training-free manner. For each class label and every available prompt, we quantify their class-specific relevance by averaging image-text relevance scores over images predicted to that class under the given prompt. These estimates are then normalized to derive class-specific weights. Evaluations on standard image classification benchmarks show that CARPRT outperforms existing class-independent reweighting methods, confirming that modeling prompt-class dependencies is crucial for effective zero-shot prediction and even broader VLM-based application settings that rely on prompt ensembling. Our code is available at https://github.com/tmlr-group/CARPRT.
CommentsAccepted at ICLR 2026