发表机构
Information Technologies Institute, CERTH(CERTH信息技术研究所)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究通过对194个视觉-语言模型的大规模实证分析,发现扩大模型规模无法缓解偏见,训练数据属性对偏见鲁棒性更关键,架构选择效果依赖基准特性。
AI 中文摘要
CLIP等视觉-语言模型(VLMs)现已成为多模态系统的基础,但人们对其在大规模应用中对虚假关联的鲁棒性仍知之甚少。我们开展了针对194个公开VLMs的首次大规模实证研究,涵盖16个模型家族、广泛的模型规模、24个训练数据集及三个评估基准:ImageNet(整体性能)、CelebA(典型单属性偏见)和UrbanCars(复杂多属性偏见)。在这些设置中,随着评估从ImageNet(斯皮尔曼相关系数ρ=0.68)转向单属性偏见基准(ρ=0.48),再到多属性偏见基准(ρ=0.05),模型规模与性能的相关性逐渐减弱。相比之下,训练数据的属性(规模与质量)与两个偏见基准的最差组准确率均呈现更一致的关系。值得注意的是,在规模相当的情况下,经过精心整理的数据集比未整理的数据集最多可提升25%的性能。最后,架构选择(如patch大小、图像分辨率)的效果高度依赖上下文,随基准的性质(包括偏见类型及其在图像中的空间分布)而变化。
英文摘要
Vision-Language Models (VLMs) such as CLIP are now foundational to multimodal systems, yet their robustness to spurious correlations remains poorly understood at scale. We present the first large-scale empirical study of 194 publicly available VLMs, including 16 model families, covering a wide range of model sizes, 24 training datasets, and three evaluation benchmarks, namely ImageNet (overall performance), CelebA (typical single-attribute bias), and UrbanCars (complex multi-attribute biases). Across these settings, the Spearman correlation between model scale and performance weakens as evaluation shifts from ImageNet ($ρ{=}0.68$) to single-attribute ($ρ{=}0.48$) and further to multi-attribute ($ρ{=}0.05$) bias benchmarks. In contrast, properties of the training data (size and quality) show more consistent relationships with worst-group accuracy across both bias benchmarks. Notably, curated datasets yield improvements of up to 25% over uncurated alternatives at a comparable scale. Finally, the effect of architectural choices (e.g., patch size, image resolution) is highly context-dependent, varying with the nature of the benchmark, including the type of bias and its spatial distribution within images.