arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2409.19375cs.LGcs.AIcs.CLcs.CVcs.HC

DOTA:视觉语言模型的分布性测试时自适应

DOTA: Distributional Test-Time Adaptation of Vision-Language Models

  • State Key Laboratory of Networking and Switching Technology(网络与交换技术国家重点实验室)
  • Beijing University of Posts and Telecommunications(北京邮电大学)
  • College of Intelligence and Computing(智能与计算学院)
  • Tianjin University(天津大学)
  • School of Computer Science and Technology(计算机科学与技术学院)
  • Harbin Institute of Technology Shenzhen(哈尔滨工业大学深圳学院)
  • Institute for Infocomm Research, A*STAR(信息通信研究机构,A*STAR)
  • Show Lab, National University of Singapore(新加坡国立大学Show实验室)
  • National University of Singapore(新加坡国立大学)

机构由 AI 辅助整理,请以论文原文为准。

Zongbo Han, Jialong Yang, Guangyu Wang, Junfan Li, Qianli Xu, Mike Zheng Shou, Changqing Zhang

更新

AI总结:

DOTA提出一种分布性测试时自适应方法,通过持续估计测试数据分布并利用贝叶斯定理计算后验概率,缓解缓存方法的灾难性遗忘,实现最先进性能。

AI中文摘要:

视觉语言基础模型(如CLIP)在广泛的任务中展现出卓越的性能。然而,当训练数据与测试数据之间存在显著分布差距时,部署这些模型可能变得不可靠,而为多样场景进行微调通常成本高昂。基于缓存的测试时适配器通过存储代表性测试样本来指导后续分类,提供了一种高效的替代方案。然而,这些方法通常采用容量有限的简单缓存管理,导致在更新过程中样本不可避免地被丢弃时出现严重的灾难性遗忘。在本文中,我们提出DOTA(分布性测试时自适应),一种简单而有效的方法来解决这一局限。关键在于,DOTA并非仅仅记忆单个测试样本,而是持续估计测试数据流的潜在分布。随后,利用贝叶斯定理,基于这些动态估计的分布计算测试时后验概率以进行自适应。这种以分布为中心的方法使模型能够持续学习并适应部署环境。大量实验验证,DOTA显著缓解了遗忘问题,并与现有方法相比实现了最先进的性能。

英文摘要:

Vision-language foundation models (VLMs), such as CLIP, exhibit remarkable performance across a wide range of tasks. However, deploying these models can be unreliable when significant distribution gaps exist between training and test data, while fine-tuning for diverse scenarios is often costly. Cache-based test-time adapters offer an efficient alternative by storing representative test samples to guide subsequent classifications. Yet, these methods typically employ naive cache management with limited capacity, leading to severe catastrophic forgetting when samples are inevitably dropped during updates. In this paper, we propose DOTA (DistributiOnal Test-time Adaptation), a simple yet effective method addressing this limitation. Crucially, instead of merely memorizing individual test samples, DOTA continuously estimates the underlying distribution of the test data stream. Test-time posterior probabilities are then computed using these dynamically estimated distributions via Bayes' theorem for adaptation. This distribution-centric approach enables the model to continually learn and adapt to the deployment environment. Extensive experiments validate that DOTA significantly mitigates forgetting and achieves state-of-the-art performance compared to existing methods.

补充信息

↑