当语言胜于图像:理解视觉-语言模型中的语言偏差
When Words Speak Louder than Images: Towards Understanding Language Bias in Vision-Language Models
浏览论文内容
中文总结 AI 辅助
本研究提出一个四阶段诊断框架,通过分析语言先验与跨模态覆盖度的演变,系统追踪并揭示视觉-语言模型中语言偏差的传播机制与成因。
中文摘要 AI 辅助
尽管视觉-语言模型(VLM)在下游应用中取得了显著进展,但它们仍然容易受到语言偏差的影响,常常优先考虑语言线索而非视觉证据,从而产生错误的预测。先前的研究提出了各种方法来理解和缓解VLM中的语言偏差,然而由于难以追踪语言偏差在黑盒VLM内部如何传播,这些研究的发现常常相互矛盾。基于单词补全任务,我们通过以下方式追踪语言偏差在VLM推理中的传播:(1)提出一个诊断框架,将推理过程分解为四个不同但相互依存的阶段,以追踪语言偏差的传播;(2)考察语言偏差背后的两个关键因素,即语言先验和跨模态覆盖度,如何在这些阶段中演变并最终导致错误预测。语言先验捕捉了VLM中语言模型组件引起的统计偏差强度,代表了语言偏差的起源,而跨模态覆盖度衡量了语言线索覆盖视觉内容的程度。通过将推理分解为四个阶段并刻画语言先验与跨模态覆盖度在这些阶段中的相互作用,我们提出了一个系统框架来追踪语言偏差在整个推理过程中的传播;并通过揭示语言先验与跨模态覆盖度之间的相互作用,揭示了语言偏差的潜在机制。
英文摘要
Despite substantial progress across downstream applications, vision-language models (VLMs) remain susceptible to language bias, often prioritizing linguistic cues over visual evidence and consequently producing incorrect predictions. Prior studies have proposed various approaches to understanding and mitigating language bias in VLMs, yet their findings often conflict due to the difficulty of tracing how language bias propagates within black-box VLMs. Building on the word completion task, we trace how language bias propagates through VLM inference by (1) proposing a diagnostic framework that decomposes the inference process into four distinct yet interdependent stages to trace the propagation of language bias; and (2) examining how two key factors underlying language bias, i.e., linguistic priors and cross-modal coverage, evolve across these stages and ultimately give rise to incorrect predictions. The linguistic prior captures the strength of statistical bias induced by the language model component of a VLM and represents the origin of language bias, whereas cross-modal coverage measures the extent to which linguistic cues cover the visual content. By decomposing inference into four stages and characterizing the interplay between linguistic priors and cross-modal coverage across these stages, we propose a systematic framework for tracing the propagation of language bias throughout the inference process; and uncover the underlying mechanism of language bias by revealing the interplay between linguistic priors and cross-modal coverage.
发表机构
- University of Waterloo(滑铁卢大学)
- Johns Hopkins University(约翰霍普金斯大学)
- University of California, Santa Barbara(加州大学圣塔芭芭拉分校)
- Nanyang Technological University(南洋理工大学)
- Indiana University(印第安纳大学)
机构由 AI 辅助整理,请以论文原文为准。