arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2405.01468cs.LGcs.AIcs.CV

理解视觉-语言模型的检索增强任务适应

Understanding Retrieval-Augmented Task Adaptation for Vision-Language Models

  • University of Wisconsin-Madison(威斯康星大学麦迪逊分校)

机构由 AI 辅助整理,请以论文原文为准。

Yifei Ming, Yixuan Li

更新

AI总结:

本文系统研究了检索增强适应中单模态与跨模态检索的作用,揭示logit集成的关键性,并给出理论支撑,为视觉-语言模型在低数据场景下的适应提供新见解。

AI中文摘要:

预训练的对比式视觉-语言模型在广泛任务上表现出色,但在微调数据集上,若类别在预训练阶段未被充分表示,模型往往表现不佳,因此需要进行适应。近期工作通过利用网络规模数据库中的样本进行检索增强适应,尤其是在低数据场景下,取得了有前景的结果。尽管在经验上取得了成功,但理解检索如何影响视觉-语言模型的适应仍是一个未解决的研究问题。本文采用反思性视角,通过系统研究来理解检索增强适应中关键组件的作用。我们揭示了单模态检索和跨模态检索的新见解,并强调了logit集成对于有效适应的关键作用。我们进一步给出了直接支持经验观察的理论基础。

英文摘要:

Pre-trained contrastive vision-language models have demonstrated remarkable performance across a wide range of tasks. However, they often struggle on fine-trained datasets with categories not adequately represented during pre-training, which makes adaptation necessary. Recent works have shown promising results by utilizing samples from web-scale databases for retrieval-augmented adaptation, especially in low-data regimes. Despite the empirical success, understanding how retrieval impacts the adaptation of vision-language models remains an open research question. In this work, we adopt a reflective perspective by presenting a systematic study to understand the roles of key components in retrieval-augmented adaptation. We unveil new insights on uni-modal and cross-modal retrieval and highlight the critical role of logit ensemble for effective adaptation. We further present theoretical underpinnings that directly support our empirical observations.

补充信息

↑