arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

从推理到适应:视觉语言模型的统一最优传输视角

From Inference to Adaptation: A Unified Optimal Transport View of Vision Language Model

Qi Yu, Zhichen Zeng, Katherine Tieu, Xiyuan Yang, Ruizhong Qiu, Yuchen Yan, Lihui Liu, Yanjun Zhao, Lingjie Chen, Jingrui He, Hanghang Tong

arXiv 2608.18339首次发表:更新:

发表机构

University of Illinois Urbana-Champaign; Amazon; Wayne State University(伊利诺伊大学厄巴纳-香槟分校; 亚马逊公司; 韦恩州立大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究提出名为algname的VLM测试时适应方法,通过Wasserstein最优传输统一VLM的推理与适应目标,实验显示其性能较最优方法提升7%且效率先进。

AI 中文摘要

视觉语言模型(Vision-language models, VLMs)已展现出出色的零样本能力,但在推理过程中对现实世界的分布偏移仍较为敏感。尽管已有大量研究致力于在测试时对VLMs进行适应,但这些方法严重依赖于推理时直接从原始嵌入相似度预测的含噪伪标签,这类伪标签在分布偏移下不可靠,会误导适应过程。为避免噪声放大,现有工作在适应过程中设计了粗粒度的代理目标,这类目标无法显式建模跨不同模态的样本级关系,导致与推理目标不匹配,进而仅能带来有限的性能提升。本研究旨在弥合VLMs的推理与适应之间脱节的目标,提出一种名为\textbf{algname}的原则性VLM测试时适应(Test-Time Adaptation, TTA)方法。对于VLM推理,我们将零样本图像分类任务表述为通过Wasserstein最优传输(Optimal Transport, OT)公式编码的跨模态对齐问题,提供样本级的鲁棒伪标签以有效适应VLMs;对于VLM适应,我们采用软标签InfoNCE损失,基于OT生成的伪标签对VLMs进行适应,利用细粒度监督通过对比学习显式建模单个图像-文本对的关系,从而在相同粒度下实现准确推理。此外,我们从理论上证明InfoNCE损失可被巧妙地重新表述为Wasserstein OT公式,进而统一VLMs的推理与适应目标以实现两者的互利。大量实验表明,我们的方法兼具有效性与效率,在达到最先进效率的同时,性能较最优方法提升了高达7%。

英文摘要

Vision-language models (VLMs) have demonstrated remarkable zero-shot capabilities yet remain sensitive to real-world distribution shifts during inference. Although significant efforts are devoted to adapting VLMs at test time, they rely heavily on noisy pseudo-labels predicted directly from raw embedding similarities during inference, which are unreliable under distribution shift and mislead the adaptation. To avoid noise amplification, existing works craft coarse-grained surrogate objectives during adaptation, which fail to explicitly model sample-level relationships across different modalities, creating objective mismatch with inference, thus leading to marginal performance improvement. In this work, we aim to bridge the detached objectives of inference and adaptation for VLMs, and propose a principled VLM TTA method called \algname. For VLM inference, we formulate the zero-shot image classification task as a cross-modal alignment problem encoded via a Wasserstein OT formulation, providing robust pseudo-labels at the sample-level to effectively adapt VLMs. For VLM adaptation, we adopt a soft-label InfoNCE loss to adapt VLMs based on the OT-induced pseudo-labels, leveraging fine-grained supervisions to explicitly model relationships of individual image-text pairs via contrastive learning, which empowers accurate inference at the same granularity. Moreover, we theoretically reveal that the InfoNCE loss can be neatly reformulated as a Wasserstein OT formulation, thereby unifying the objectives of the inference and adaptation of VLMs to achieve their mutual benefits. Extensive experiments demonstrate the effectiveness and efficiency of our methods, outperforming the best-performing methods by up to 7% with state-of-the-art efficiency.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑