Do Vision-Language Models See or Guess? Measuring and Reducing Textual-Prior Reliance with a Phrasing-Controlled Benchmark
视觉语言模型是看见还是猜测?通过措辞控制基准衡量和减少文本先验依赖
机构 * Lossfunk ; Indian Institute of Technology Roorkee(印度理工学院罗尔基分校) ; Raeth AI
AI总结 本文构建了540张图像的基准,通过为同一图像生成四种措辞变体,衡量视觉语言模型对文本先验的依赖,发现所有模型在最难变体上性能下降,开放模型下降最严重,并通过无图像消融等分析证实了真正的图像依赖。
Comments 17 pages, 7 figures, Submitted to EMNLP 2026