归纳视觉逻辑用于视觉语言模型的小样本分布外适应
Inductive Visual Logic for Few-Shot Out-Of-Distribution Adaptation in VLMs
浏览论文内容
中文总结 AI 辅助
针对视觉语言模型在远距离分布外任务上的判别失败,提出无需训练的归纳视觉逻辑框架,利用模型保留的描述能力构建分类知识,在多个基准上取得最优准确率。
中文摘要 AI 辅助
生成式视觉语言模型(VLMs),如Qwen-VL和LLaVA,在与预训练分布重叠的任务上展现出强大的零样本性能,但在需要从未学习过的判别性特征的专业领域上失败,我们将这一领域称为远距离分布外(OOD)。标准适应方法无法克服这种表征缺失,因为它们是在编码器现有的特征空间内操作。然而,即使判别能力崩溃,VLMs仍保留强大的描述能力:一个无法对医学扫描进行分类的模型仍能清晰阐述其视觉模式。利用这种不对称性,我们引入了归纳视觉逻辑(IVL),一个无需训练的框架,从模型存活的描述能力中构建分类知识。IVL通过双模式提示从少量支持图像中提取视觉特征,将语义描述与原始视觉观察相结合,并将其组织为每类特征字典。在推理时,分层过滤识别空间定位的特征证据以进行分类。在多个远距离OOD基准测试中,IVL在两种VLM骨干下实现了最高的总体准确率,同时产生可解释的、特征可追溯的预测。
英文摘要
Generative vision-language models (VLMs) such as Qwen-VL and LLaVA achieve strong zero-shot performance on tasks overlapping with their pretraining distribution, yet fail on specialized domains where the required discriminative features were never learned, a regime we term distant out-of-distribution (OOD). Standard adaptation methods cannot overcome this representational absence because they operate within the encoder's existing feature space. However, VLMs retain a robust descriptive capacity even when discrimination collapses: a model that cannot classify a medical scan can still articulate its visual patterns. Exploiting this asymmetry, we introduce Inductive Visual Logic (IVL), a training-free framework that constructs classification knowledge from the model's surviving descriptive ability. IVL extracts visual traits from few-shot support images through dual-mode prompting, combining semantic descriptions with primitive visual observations, and organizes them into per-class trait dictionaries. At inference, hierarchical filtering identifies spatially grounded trait evidence for classification. Across multiple distant-OOD benchmarks, IVL achieves the highest aggregate accuracy under two VLM backbones while producing interpretable, trait-traceable predictions.
发表机构
- National Tsing Hua University(国立清华大学)
- National Taiwan University(国立台湾大学)
机构由 AI 辅助整理,请以论文原文为准。