发表机构
National Institute of Advanced Industrial Science and Technology (AIST); University of Tsukuba; University of Technology Nuremberg; University of Oxford(日本产业技术综合研究所; 筑波大学; 纽伦堡工业大学; 牛津大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究发现VLMs存在颜色偏差,通过引入Stealth Visual Prompts控制文本视觉风格,发现颜色、对比度会影响VLMs的情感预测和VQA输出,其行为与视觉编码器潜在表示变化相关。
AI 中文摘要
视觉语言模型(Vision Language Models, VLMs)正越来越多地被用于招聘支持、推荐等工业决策系统,这促使人们仔细分析VLMs如何处理视觉和文本信息。在本研究中,我们探究VLMs如何解释以图像形式呈现的文本,并研究视觉风格偏差的影响。为此,我们引入了隐形视觉提示(Stealth Visual Prompts),该提示可在保留语义内容的同时,细微改变文本的视觉风格,如颜色和对比度。我们利用这些提示系统地控制文本中词语的视觉风格,并测量其对VLMs分析的影响。我们还进一步分析此类视觉扰动如何影响视觉编码器的潜在表示。实验中,我们观察到将积极词汇染成绿色会持续将情感预测推向积极方向,导致VLMs往往无法正确考虑文本中存在的消极词汇。我们的分析表明,这种行为与颜色变化引起的视觉编码器潜在表示的变化相关。此外,我们还发现降低文本-背景对比度会增加对视觉显著线索的依赖,导致视觉问答(Visual Question Answering, VQA)输出错误更多。这些结果表明,渲染文本的视觉风格可以引导VLMs的解释,且这种引导方式与人类的语义理解存在偏差。项目页面:this https URL
英文摘要
Vision language models (VLMs) are increasingly used in industrial decision-making systems, such as recruitment support and recommendation. This motivates careful analysis of how VLMs process visual and textual information. In this work, we study how VLMs interpret text rendered as an image, and investigate the influence of visual styling biases. To this end, we introduce Stealth Visual Prompts, which subtly change visual styling of text, such as color and contrast, while preserving semantic content. Using these prompts, we systematically control the visual styling of words in text and measure their impact on the analysis performed by VLMs. We further analyze how such visual perturbations affect the latent representations of the vision encoder. From our experiments, we observed that coloring positive words in green consistently shifts sentiment predictions toward a positive direction. As a result, VLMs often fail to properly account for negative words present in the text. Our analysis suggests that this behavior is correlated with changes in the latent representations of the vision encoder induced by color variations. In addition, we show that reducing text--background contrast increases reliance on visually salient cues and leads to more incorrect Visual Question Answering (VQA) outputs. These results suggest that the visual styling of rendered text can guide VLMs' interpretation in ways that diverge from human semantic understanding. Project page: https://github.com/KohsukeIde/color-bias-vlm
Comments15 pages. Accepted to ICPR 2026
Journal refIn: Pattern Recognition. ICPR 2026. Lecture Notes in Computer Science. Springer, Cham, pp. 261-275
DOI:10.1007/978-3-032-31583-0_18