读对了,答错了:视觉配置变化如何影响视觉语言模型中的证据使用
Reading Right, Answering Wrong: How Visual Configuration Changes Affect Evidence Use in VLMs
- Northeastern University(东北大学)
- NiuTrans Research(纽创思研究)
机构由 AI 辅助整理,请以论文原文为准。
中文总结 AI 辅助
本研究揭示视觉配置变化(如图像分块与标记排列)导致视觉语言模型在问答中不稳定,即使模型仍能读取正确信息,其使用能力却减弱;通过字段线索与转录引导,可纠正97.2%的相关错误。
中文摘要 AI 辅助
视觉语言模型(VLMs)在视觉问答等任务上取得了强劲的性能,然而,微小的图像尺寸调整就能将正确答案转变为错误。我们研究了视觉配置的变化(如图像分块和标记排列)是否导致了这种不稳定性。在七个检查点和四个基准测试中,同样微小的尺寸调整在切换配置时会引起更多的正确性翻转。令人惊讶的是,在这些案例中,超过一半的情况下,模型回答问题错误,但当被告知要读取什么时,它们仍能读取正确的答案。此外,对LLaVA-NeXT的注意力干预表明,配置变化会削弱在回答过程中对可读信息的使用。因此,我们利用字段线索和模型自身的转录来引导模型。在标注辅助下,这些引导形式共同纠正了97.2%的带有可读信息的错误。这些发现表明,配置变化会影响模型如何使用它们仍能读取的信息。
英文摘要
Vision-language models (VLMs) have achieved strong performance on tasks such as visual question answering, yet small image resizes can turn correct answers into errors. We investigate whether changes in visual configuration, such as image tiling and token arrangement, contribute to this instability. Across seven checkpoints and four benchmarks, equally small resizes cause more correctness flips when they switch configurations. Surprisingly, in over half of these cases, models answer the question incorrectly but can still read the correct answer when told what to read. Furthermore, attention interventions in LLaVA-NeXT suggest that configuration changes can weaken the use of readable information during answering. We therefore guide models using field cues and their own transcriptions. With annotation assistance, these forms of guidance together correct 97.2% of errors with readable information. These findings show that configuration changes can affect how models use information they can still read.