通过结构化提示重参数化重新编程视觉-语言模型
Reprogramming Vision-Language Models via Structured Prompt Reparameterization
- The University of Melbourne(墨尔本大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
提出结构化提示重参数化方法RVP,通过类内聚合和类间残差校正,在保持CLIP主干冻结下实现高效视觉重编程,在11个少样本基准上优于现有方法。
AI中文摘要:
视觉重编程通过修改预训练模型的输入和输出接口而保持主干网络固定,从而使其适应下游任务。在视觉-语言模型中,现有方法主要依赖于类内提示聚合,并未显式建模类间关系。然而,细粒度类别在文本嵌入空间中往往表现出高度重叠的属性描述和强烈的类间相关性,其中判别性线索存在于微妙的低方差分量中。我们提出了重参数化类间视觉重编程(RVP),这是一种结构化框架,在每个类内聚合多个文本提示,并在类间应用残差校正。我们还表明,基于CLIP的视觉重编程,若采用与输入无关的线性输出聚合,可以表示为从冻结图像嵌入到下游逻辑的线性映射,并利用这一观点设计了一种结构化重参数化,以建模共享语义组件和类特定差异。RVP仅使用单个视觉提示,并可在推理时重参数化为冻结主干网络后接线性分类器,几乎不增加计算开销。在11个少样本分类基准和四个CLIP主干网络上,RVP一致优于先前的视觉重编程方法,且推理效率相当或更好。
英文摘要:
Visual reprogramming adapts pretrained models to downstream tasks by modifying their input and output interfaces while keeping the backbone fixed. In vision-language models, existing methods mainly rely on intra-class prompt aggregation and do not explicitly model relationships among classes. However, fine-grained categories often exhibit highly overlapping attribute descriptions and strong inter-class correlation in the text embedding space, where discriminative cues lie in subtle low-variance components. We propose Reparameterized Inter-Class Visual Reprogramming (RVP), a structured framework that aggregates multiple text prompts within each class and applies residual correction across classes. We also show that CLIP-based visual reprogramming with input-independent linear output aggregation can be expressed as a linear mapping from frozen image embeddings to downstream logits, and use this view to design a structured reparameterization that models shared semantic components and class-specific differences. RVP uses only a single visual prompt and can be reparameterized at inference into a frozen backbone followed by a linear classifier, incurring nearly zero computational overhead. Across 11 few-shot classification benchmarks and four CLIP backbones, RVP consistently improves over prior visual reprogramming methods with comparable or better inference efficiency.