arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

域重定心与置信度加权先验校准用于视觉-语言模型

Domain Recentering and Confidence-Weighted Prior Calibration for Vision-Language Models

Youngeun Seol, Jimin Shin, Heeseo Yoon, Uiwon Hwang

arXiv 2609.29358首次发表:更新:

发表机构

Ewha Womans University(梨花女子大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对视觉-语言模型在分布偏移下性能下降的问题,提出免训练的域重定心与置信度校准方法,通过高斯混合后验加权和置信度加权先验校正,提升跨域零样本分类准确率。

AI 中文摘要

诸如CLIP之类的视觉-语言模型实现了强大的零样本分类,然而在分布偏移下,视觉嵌入会偏离固定的文本嵌入。免训练校准避免了提示学习中的逐样本优化,但先验特征校准会给每个图像带来一个硬聚类的全部偏差。我们提出了带有置信度校准的域重定心(DRC),这是一种免训练方法,从一组未标记的目标图像中调整CLIP。DRC拟合一次高斯混合,并从每个嵌入中减去组件均值的后验加权平均值。然后,它通过对数先验校正去除残余的类别偏好,从置信度加权预测中估计先验。在比较的方法中,DRC在跨域数据集上取得了最高的平均准确率,使用ViT-B/16和ResNet-50分别超过零样本CLIP 4.13和5.07个百分点,且在ImageNet分布偏移下,相对于CLIP的提升仍然保持。

英文摘要

Vision-language models such as CLIP achieve strong zero-shot classification, yet under distribution shift, visual embeddings drift from fixed text embeddings. Training-free calibration avoids the per-sample optimization of prompt learning, but prior feature calibration gives each image the full bias of one hard cluster. We propose Domain Recentering with Confidence Calibration (DRC), a training-free method adapting CLIP from a set of unlabeled target images. DRC fits a Gaussian mixture once and subtracts from each embedding a posterior-weighted average of component means. It then removes residual class preference with a log-prior correction, estimating the prior from confidence-weighted predictions. Among compared methods, DRC achieves the highest average accuracy on cross-domain datasets, exceeding zero-shot CLIP by 4.13 and 5.07 points with ViT-B/16 and ResNet-50, with gains over CLIP also holding under ImageNet distribution shifts.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑