arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

GATE:用于视觉语言模型无训练测试时自适应的可靠性门控高斯证据融合

GATE: Reliability-Gated Gaussian Evidence Fusion for Training-Free Test-Time Adaptation of Vision-Language Models

Pedram MohajerAnsari, Amir Salarpour, Run Wang, Mert D. Pesé

arXiv 2608.29395首次发表:更新:

发表机构

Clemson University(克莱姆森大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究提出无训练的两阶段直推式TTA框架GATE,通过可靠性门控融合文本与图像高斯证据,在多数据集及模型上显著提升视觉语言模型的零样本测试时自适应性能。

AI 中文摘要

CLIP和SigLIP等视觉语言模型(VLM)具备强大的零样本识别能力,但当部署到与预训练分布不同的目标数据时,其预测性能会下降。测试时自适应(TTA)是无需源数据或目标标签即可提升鲁棒性的实用方法,然而现有方法往往仅依赖提示侧自适应或图像侧目标证据。本研究提出GATE,这是一种无训练的两阶段直推式测试时自适应框架,利用未标记目标集,同时保持图像编码器、文本编码器和提示参数完全冻结。GATE不使用单个原型表示每个类别,而是在共享的视觉语言特征空间中构建两个互补的高斯证据源:从多个语言描述估计得到的文本高斯,以及从可靠未标记目标样本估计得到的图像高斯。类别级可靠性门控控制图像衍生伪证据的影响,分数级广义专家乘积融合对原始零样本logits产生归一化残差修正。在细粒度识别数据集、ImageNet系列分布偏移、多个CLIP主干及SigLIP-B/16上的实验显示,GATE在每个基准/主干组中均取得最佳平均准确率,平均提升零样本性能5.41个百分点,且比最强的非GATE基线高出1.94个百分点,证明了可靠性门控分布证据对冻结VLM自适应的益处。

英文摘要

Vision-language models such as CLIP and SigLIP provide strong zero-shot recognition, but their predictions can degrade when deployed on target data that differ from the pretraining distribution. Test-time adaptation offers a practical way to improve robustness without source data or target labels, yet existing methods often rely on either prompt-side adaptation or image-side target evidence alone. In this work, we introduce GATE, a training-free two-pass transductive test-time adaptation framework that uses the unlabeled target set while keeping the image encoder, text encoder, and prompt parameters fully frozen. Instead of representing each class with a single prototype, GATE builds two complementary Gaussian sources of evidence in the shared vision-language feature space: a text Gaussian estimated from multiple language descriptions and an image Gaussian estimated from reliable unlabeled target samples. A class-wise reliability gate controls the influence of image-derived pseudo-evidence, and a score-level generalized Product-of-Experts fusion produces a normalized residual correction to the original zero-shot logits. Across fine-grained recognition datasets, ImageNet-family distribution shifts, multiple CLIP backbones, and SigLIP-B/16, GATE achieves the best average accuracy in every benchmark/backbone group. It improves zero-shot performance by an average of 5.41 points and outperforms the strongest non-GATE baseline by 1.94 points, demonstrating the benefit of reliability-gated distributional evidence for frozen VLM adaptation.

CommentsAccepted to the British Machine Vision Conference (BMVC 2026)

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑