arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2401.12425cs.CVcs.CLcs.LG

视觉语言模型中被忽视的长尾概念

The Neglected Tails in Vision-Language Models

  • Texas A&M University(德克萨斯农工大学)
  • Carnegie Mellon University(卡内基梅隆大学)
  • Zhejiang Lab(之江实验室)
  • University of Macau(澳门大学)

机构由 AI 辅助整理,请以论文原文为准。

Shubham Parashar, Zhiqiu Lin, Tian Liu, Xiangjue Dong, Yanan Li, Deva Ramanan, James Caverlee, Shu Kong

更新

AI总结:

针对视觉语言模型在长尾概念上性能不佳的问题,提出REAL方法,利用LLM统计同义词频率并检索均衡数据训练线性分类器,以400倍更少存储和10000倍更少训练时间超越零样本SOTA。

AI中文摘要:

视觉语言模型(VLMs)在零样本识别方面表现出色,但其在不同视觉概念上的性能差异很大。例如,尽管CLIP在ImageNet上取得了令人印象深刻的准确率(60-80%),但对于夜蛇等十多个概念,其性能却降至10%以下,这可能是由于这些概念在预训练数据中的出现频率较低。然而,衡量VLM大规模数据集中概念的频率颇具挑战。我们通过使用大型语言模型(LLMs)来统计包含这些概念同义词的预训练文本数量,从而解决了这一问题。我们的分析证实,诸如LAION之类的流行数据集呈现出长尾概念分布,导致VLM的性能存在偏差。我们还发现,VLM的下游应用,包括视觉聊天机器人(如GPT-4V)和文本到图像模型(如Stable Diffusion),往往无法识别或生成我们方法所识别的稀有概念图像。为了缓解零样本VLM的性能不均衡问题,我们提出了检索增强学习(REAL)。首先,REAL不使用原始类别名称提示VLM,而是使用预训练文本中这些类别最频繁出现的同义词。这一简单改动已在九个基准数据集上超越了昂贵的人工设计和LLM增强提示。其次,REAL使用通过概念同义词检索到的小型但均衡的预训练数据子集训练线性分类器。REAL超越了之前的零样本最先进水平,同时使用的存储空间减少了400倍,训练时间减少了10,000倍!

英文摘要:

Vision-language models (VLMs) excel in zero-shot recognition but their performance varies greatly across different visual concepts. For example, although CLIP achieves impressive accuracy on ImageNet (60-80%), its performance drops below 10% for more than ten concepts like night snake, presumably due to their limited presence in the pretraining data. However, measuring the frequency of concepts in VLMs' large-scale datasets is challenging. We address this by using large language models (LLMs) to count the number of pretraining texts that contain synonyms of these concepts. Our analysis confirms that popular datasets, such as LAION, exhibit a long-tailed concept distribution, yielding biased performance in VLMs. We also find that downstream applications of VLMs, including visual chatbots (e.g., GPT-4V) and text-to-image models (e.g., Stable Diffusion), often fail to recognize or generate images of rare concepts identified by our method. To mitigate the imbalanced performance of zero-shot VLMs, we propose REtrieval-Augmented Learning (REAL). First, instead of prompting VLMs using the original class names, REAL uses their most frequent synonyms found in pretraining texts. This simple change already outperforms costly human-engineered and LLM-enriched prompts over nine benchmark datasets. Second, REAL trains a linear classifier on a small yet balanced set of pretraining data retrieved using concept synonyms. REAL surpasses the previous zero-shot SOTA, using 400x less storage and 10,000x less training time!

补充信息

↑