arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.28967cs.CVcs.LG

用于高效提示调优的视觉分布锚定

Visual Distribution Anchoring for Efficient Prompt Tuning

  • University of Minnesota(明尼苏达大学)

机构由 AI 辅助整理,请以论文原文为准。

Pouya Parsa, Raoof Zare Moayedi, Seongjin Choi

AI总结:

本文提出无需训练的VDA框架,用未标记目标池的类级视觉原型增强冻结语义分类器,在10次ImageNet到目标域迁移中显著提升多个视觉-语言模型性能,且与各类分类器互补。

AI中文摘要:

提示调优使用少量可训练参数适配视觉-语言模型,但现有方法在效率与适配性间存在权衡:静态文本提示可能过拟合源类,图像条件提示增加单实例计算量,多模态调优则修改视觉分支。本文提出VDA(Visual Distribution Anchoring,视觉分布锚定),一种无需训练的目标适配框架,它通过从未标记目标池中离线估计的类级视觉原型,增强冻结的语义分类器。本文首先探究原型能否从类名合成:文本到质心映射器可重建保留的源原型,但在数据集偏移下失效,因为类名仅指定语义身份而非目标域外观;oracle分析证实真实目标原型具有高判别力。因此VDA使用冻结的语义和域模板分类器,将未标记目标图像划分为与类相关的组,经置信度排序的图像特征形成归一化原型,再用一个全局权重与语义分类器融合。该适配无需目标标签、目标侧优化、均匀类先验假设、迭代优化或测试查询访问,可生成固定且可缓存的分类器。控制实验表明,类特定分区是性能提升的关键,且视觉局部伪标签错误虽类不正确但仍具可用性。在10次ImageNet到目标域的迁移中,该冻结设计分别将零样本CLIP、TCP和MaPLe提升3.22、3.39和3.35个点,在所有设置中改进了9个目标;其视觉修正还将无泄漏的PromptKD提升2.79个点,可与零样本、源提示、多模态提示及目标蒸馏分类器互补。

英文摘要:

Prompt tuning adapts vision--language models with few trainable parameters, but existing approaches trade off efficiency and adaptation: static textual prompts can overfit source classes, image-conditioned prompts add per-instance computation, and multimodal tuning modifies the visual branch. We propose VDA (Visual Distribution Anchoring), a training-free target adaptation framework that augments a frozen semantic classifier with class-level visual prototypes estimated offline from an unlabeled target pool. We first ask whether prototypes can be synthesized from class names. A text-to-centroid mapper reconstructs held-out source prototypes but fails under dataset shift because class names specify semantic identity, not target-domain appearance. An oracle analysis confirms that true target prototypes are highly discriminative. VDA therefore uses frozen semantic and domain-template classifiers to partition unlabeled target images into class-correlated groups. Confidence-ranked image features form normalized prototypes, fused with the semantic classifier using one global weight. Adaptation requires no target labels, target-side optimization, uniform class-prior assumption, iterative refinement, or test-query access, and yields a fixed, cacheable classifier. Controlled experiments show that class-specific partitioning drives gains and that visually local pseudo-label errors can remain useful despite being class-incorrect. Across ten ImageNet-to-target transfers, the same frozen design improves zero-shot CLIP, TCP, and MaPLe by 3.22, 3.39, and 3.35 points, respectively, improving nine of ten targets in every setting. Its visual correction further improves leakage-free PromptKD by 2.79 points, complementing zero-shot, source-prompted, multimodal-prompted, and target-distilled classifiers.

补充信息

↑