混合比例并非性能所在:诊断视觉-语言模型少样本适配中的原型混合
The Blending Ratio Is Not Where the Performance Is: Diagnosing Prototype Blending for Few-Shot Adaptation of Vision-Language Models
浏览论文内容
中文总结 AI 辅助
本文针对视觉-语言模型少样本适配的原型混合方法,发现其最优混合比例可仅通过支持集估计,且无验证的线性探针性能优于最优调参混合,性能上限由模型本身而非超参数决定。
中文摘要 AI 辅助
许多视觉-语言模型的少样本适配方法,通过零样本文本原型与K个标记图像特征均值的凸组合进行分类,且通常在保留标签(常为测试集本身)上调整单一混合比例。本文探究该方法的偏差-方差合理性所引发的问题:合适的比例是什么?能否在无验证数据时估计?性能是否由该比例决定?首先,使原型均方误差最小的比例有闭式解,其支持集插件恰好是向文本原型收缩的正部James-Stein系数。在4800个单元(10个数据集、5个包括SigLIP在内的骨干网络、5个样本数、5个随机种子、4个提示层级)中,该理论最优比例对错误量的估计不可靠:在其有定义的950个主层级单元中,比测试集最优比例低8.5个点,且趋近于1,丢弃文本先验转而采用最近类均值分类器,因为其视为偏差的78%文本-图像原型距离是与类无关的偏移,可通过argmax大致抵消。本文通过反事实证明该机制,并将其损害份额上限限定为26%。其次,仅在支持集上留一法设置的比例,与最优混合比例的差距在0.9个点内,因此无需验证数据即可估计。第三,无验证的线性探针甚至优于最优调整的混合:CLAP平均提升1.9个点,LP++平均提升1.5个点,且当K≥4时,所有4个无验证基线均优于最优比例,线性探针的提升幅度不包含零。这些结果表明,模型类的性能上限由模型本身而非超参数决定:比例可近乎最优地免费设置,但性能仍不取决于此。代码、缓存特征、各单元记录:this https URL
英文摘要
Many few-shot adaptation methods for vision-language models classify with a convex combination of the zero-shot text prototype and the mean of the K labelled image features, with a single blending ratio routinely tuned on held-out labels, often on the test set itself. We ask what the family's own bias-variance justification invites: what is the right ratio, can it be estimated without validation data, and is finding it where the performance is? First, the ratio minimising prototype mean-squared error has a closed form whose support-set plug-in is exactly a positive-part James-Stein coefficient shrinking towards the text prototype. Across 4,800 cells (ten datasets, five backbones including SigLIP, five shot counts, five seeds, four prompt tiers) this theoretically optimal ratio is a reliable estimate of the wrong quantity: on the 950 primary-tier cells where it is defined it trails a test-set-oracle ratio by 8.5 points. It saturates near 1, discarding the text prior for a nearest-class-mean classifier, because 78% of the text-image prototype distance it treats as bias is a class-independent offset that the arg max largely cancels. We prove the mechanism and bound its share of the damage at 26% by a counterfactual. Second, leave-one-out on the support set alone sets a ratio landing within 0.9 points of the oracle blend, so it is estimable without validation data. Third, validation-free linear probes beat even the oracle-tuned blend: CLAP by +1.9 points and LP++ by +1.5 on average, and at K >= 4 all four validation-free baselines sit above the oracle, the linear probes by margins excluding zero. These results locate the ceiling in the model class, not the hyperparameter: the ratio can be set near-optimally for free, and it is still not where the performance is. Code, cached features, per-cell records: https://huggingface.co/datasets/Liangzhi-Li/clipbench-blending
发表机构
- Qufu Normal University(曲阜师范大学)
- Climind
- The University of Osaka(大阪大学)
- Institute of Scientific and Industrial Research (SANKEN)(产业科学研究所(SANKEN))
- Institute of High Performance Computing (IHPC), A*STAR(新加坡科技研究局高性能计算研究所(IHPC))
- Standard Chartered Bank(渣打银行)
- Hainan University(海南大学)
机构由 AI 辅助整理,请以论文原文为准。