发表机构
NIT Silchar; MBZUAI; BML Munjal University; KAUST(锡尔恰尔国立理工学院; 穆罕默德·本·扎耶德人工智能大学; BML穆贾尔大学; 阿卜杜拉国王科技大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究提出Layered-VQA基准,发现视觉语言模型在场景碎片化时性能显著下降,表明证据的组合方式比证据本身更关键。
AI 中文摘要
视觉语言模型(VLM)越来越多地基于随时间被裁剪、分割、检索或逐步揭示的视觉证据进行推理。然而,大多数VQA基准测试都是同时呈现完整图像和问题。我们探究当相同信息被碎片化时,模型会丢失什么。我们引入了Layered-VQA,包含93个场景和300个问题。每个图像被分解为有序的RGBA图层,这些图层能精确重组原始场景,每个问题都标注了支持图层、最小充分图层和干扰图层。我们评估了从3B到32B参数的十一个开放权重VLM和两个专有模型,规模为187,200次对话,由174万次开放模型交叉判断进行评分。我们发现三个一致的失败。组合损失:碎片化问题影响较小,但碎片化场景显著降低准确性;重组相同图层在很大程度上恢复了性能。预言机反转:即使预言机选择的充分证据也可能表现不如完整场景。接地损失:随着所需证据增多,接地(grounding)能力下降的速度远快于答案准确性。这些结果共同表明,拥有正确的视觉证据是不够的。证据如何组合和呈现决定了模型能否使用并接地这些证据。正确的证据是不够的:VLM需要它来源的场景。
英文摘要
Vision-language models (VLMs) increasingly reason over visual evidence that is cropped, segmented, retrieved, or revealed over time. Yet most VQA benchmarks present the complete image and question at once. We ask what models lose when the same information is fragmented. We introduce Layered-VQA, with 93 scenes and 300 questions. Each image is decomposed into ordered RGBA layers that exactly recompose the original scene, and each question is annotated with supporting, minimal-sufficient, and distractor layers. We evaluate eleven open-weight VLMs from 3B to 32B parameters and two proprietary models with a scale of 187,200 conversations, graded by 1.74M open-model cross-judgments. We find three consistent failures. Loss in Composition: fragmenting the question has a small effect, but fragmenting the scene substantially reduces accuracy; recomposing the same layers largely restores performance. Oracle Inversion: even oracle-selected sufficient evidence can perform worse than the complete scene. Loss in Grounding: as more evidence is required, grounding degrades much faster than answer accuracy. Together, these results show that having the right visual evidence is not enough. How that evidence is composed and presented determines whether models can use and ground it. The right evidence is not enough: VLMs need the scene it came from.
Comments33 pages, 10 figures, 11 tables. Code: https://github.com/lost-in-layers/composition-not-conversation . Dataset: https://huggingface.co/lost-in-layer/Layered-VQA