RADC:面向视觉-语言测试时自适应的风险感知双缓存
RADC: Risk-Aware Dual Caching for Vision-Language Test-Time Adaptation
- Tsinghua Shenzhen International Graduate School, Tsinghua University(清华大学深圳国际研究生院)
- Pengcheng Laboratory(鹏城实验室)
- School of Intelligent Systems Engineering, Sun Yat-sen University(中山大学智能工程学院)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
RADC通过语义前景缓存和高斯风险准入机制,解决视觉-语言测试时自适应中背景偏差与缓存准入不可靠问题,在跨域和分布外基准上取得最先进性能。
AI中文摘要:
基于缓存的测试时自适应(TTA)方法在视觉-语言模型中常因全局表示中的背景偏差以及表示变化下不可靠的基于熵的缓存准入而受到阻碍。为解决这些局限,我们提出RADC,通过可靠的双缓存机制增强原型学习。RADC引入一个语义前景缓存,从CLIP表示中聚合类别一致的空间证据,生成与全局缓存互补的前景原型,同时减轻背景干扰。为可靠地管理这两个缓存,高斯风险准入将多视图表示建模为对角高斯分布,并联合考虑类别分离度和特征不确定性,以优先选择可靠的缓存候选。RADC将零样本logits与互补的全局和前景缓存预测相集成,实现稳健推理。在跨域和分布外基准上的大量实验表明,其性能持续达到最先进水平。
英文摘要:
Cache-based test-time adaptation (TTA) for vision-language models is often hindered by background bias in global representations and unreliable entropy-based cache admission under representation variations. To address these limitations, we propose RADC, which enhances prototype learning through reliable dual caching. RADC introduces a Semantic Foreground Cache that aggregates category-consistent spatial evidence from CLIP representations, yielding foreground prototypes that complement the global cache while mitigating background interference. To reliably manage both caches, Gaussian Risk Admission models multi-view representations as diagonal Gaussian distributions and jointly considers class separation and feature uncertainty to prioritize reliable cache candidates. RADC integrates zero-shot logits with complementary global- and foreground-cache predictions for robust inference. Extensive experiments on cross-domain and out-of-distribution benchmarks demonstrate consistent state-of-the-art performance.