arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

PRiSM:少样本视觉语言模型的原型正则化

PRiSM: Prototype Regularization for Few-Shot VLMs

Ghassen Baklouti, Omprakash Chakraborty, Jose Dolz, Ismail Ben Ayed

arXiv 2607.17820首次发表:更新:

发表机构

ÉTS Montreal(蒙特利尔高等商学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究针对视觉语言模型少样本适应方法,质疑现有基准假设,引入新基准。提出PRiSM类原型正则化方法,优化多术语损失,结合有效优化器,提升性能,尤其在处理类不平衡和大量类时效果显著。

AI 中文摘要

无训练的少样本适应方法在视觉语言模型(VLMs)中备受关注。然而,当前基准依赖对适应数据统计的强假设,如类平衡。我们质疑这些简化假设,通过狄利克雷采样引入更现实的基准,改变少样本任务中的类平衡水平和有效类数。令人惊讶的是,在我们的设置下,现有方法性能大幅下降,尤其是标记样本增加时。为缓解此问题,我们引入PRiSM,一种可作为即插即用模块的类原型正则化方法,显著提升性能。我们的方法优化了一种新颖的多术语损失,包括最大化类间成对距离的正则化器,以及促进支持特征对齐和与基线原型保真度的附加项。此外,我们为目标引入了一种有效且计算高效的块Majorize-Minimize优化器。具体而言,我们推导出一个有效的块wise Lipschitz常数,可通过格什戈林圆定理有效计算。广泛实验表明,PRiSM改进了多个无训练基线,在处理严重类不平衡和大量类时收益显著。

英文摘要

Training-free few-shot adaptation methods have gained significant attention recently in the context of Vision-language Models (VLMs). Yet, current benchmarks rely on strong assumptions about the statistics of the adaptation data, e.g., class balance. We question these simplifying assumptions and introduce a more realistic benchmark that varies both the levels of class balance and the effective number of classes in few-shot tasks via Dirichlet sampling. Surprisingly, under our setting, we observe substantial drops in the performances of state-of-the-art methods, more so when the number of labeled samples increases. To mitigate this, we introduce PRiSM, a class-prototype regularization that can be deployed as a plug and play module on top of any existing baseline method, significantly improving performances. Our method optimizes a novel multi-term loss, which includes a regularizer maximizing inter-class pairwise distances, along with additional terms promoting support-feature alignment and fidelity to the baseline prototypes. Furthermore, we introduce an effective and computationally efficient block Majorize-Minimize optimizer for our objective. More specifically, we derive a valid blockwise Lipschitz constant (i.e., a bound on the Hessian's spectral norm), which can be computed efficiently via the Gershgorin circle theorem. Extensive experiments show that PRiSM improves several training-free baselines, with large gains when dealing with severe class imbalance and high numbers of classes.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑