发表机构
Xidian University(西安电子科技大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
G2D是一种无需训练的零样本图像分类框架,通过生成式VLM验证CLIP检索的候选,在8个基准上平均准确率达68.85%,优于CLIP及独立生成式模型,还可迁移至其他模型。
AI 中文摘要
零样本分类需要高效的标签检索和细粒度视觉推理,但判别式和生成式视觉语言模型(VLM)在互补性方面存在不足。当CLIP的Top-1预测错误时,正确标签通常仍在其Top-K候选列表中,因此消歧而非召回成为关键。然而,生成式模型受限于庞大的标签空间和无约束的推理。这种互补性促使我们将宽泛的候选检索与细粒度的、基于图像的推理相分离。我们提出G2D,这是一个无需训练的框架,它使用生成式VLM对CLIP检索到的候选进行验证,其中标签名称和CLIP概率为解决视觉相似的候选提供结构化先验。置信度路由、熵自适应候选大小调整和前缀树约束解码将生成式推理聚焦于不确定样本,并确保在测试时为每个输入提供一个有效输出。在8个基准测试中,G2D的平均准确率达到68.85%,而CLIP为59.35%,独立生成式模型为63.11%。在7种生成器配置下,候选集验证将平均准确率提高了1.08至27.42个百分点。G2D还可迁移至DCLIP、WaffleCLIP和CuPL,为判别式提议与生成式视觉推理之间提供实用接口。
英文摘要
Zero-shot classification needs efficient label retrieval and fine-grained visual reasoning, yet discriminative and generative vision-language models fail in complementary ways.When CLIP's top-1 prediction is wrong, the correct label often remains in its top-$K$ shortlist, making disambiguation rather than recall the key challenge.Standalone generative models, however, are hindered by large label spaces and unconstrained outputs.This complementarity motivates separating broad candidate retrieval from fine-grained, image-grounded verification.We propose G2D, a training-free framework that uses a generative VLM to verify CLIP-retrieved candidates against the image.Candidate names and CLIP probabilities provide a structured prior for resolving visually similar classes.Fixed confidence routing, entropy-adaptive candidate sizing, and trie-constrained decoding focus generative reasoning on uncertain samples and ensure one valid output for each input at test time.Across eight benchmarks, G2D achieves 68.85% average accuracy, versus 59.35% for CLIP and 63.11% for the standalone VLM.Across seven generator configurations, candidate-set verification improves average accuracy by 1.08--27.42 percentage points.G2D also transfers to DCLIP, WaffleCLIP, and CuPL, supporting a practical interface between discriminative proposal and generative visual reasoning. Code: https://github.com/Harzva/G2D
CommentsAccepted at ACM MM 2026. 10 pages, 5 figures
Journal refProceedings of the 34th ACM International Conference on Multimedia (MM '26), 2026