发表机构
Manipal University Jaipur(马尼帕尔大学斋浦尔分校)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
GroundBench通过分层反事实条件分离视觉基础与类别-动作关联,定位VLM可供性失败,发现仅提供部件类别即可提升准确率,而视觉信息并非必要。
AI 中文摘要
一项配套评估发现,在操作提示中命名目标部件可将八个视觉语言模型的动作准确率提高0.32-0.63,且在部件被命名之前,没有模型能超越恒定基线。然而,命名部件提供了真实系统必须推断的信息,混淆了视觉基础、机械推理和类别到动作的关联。我们引入了GroundBench,一个诊断基准,通过六个分支合并条件分离这些解释,每个条件添加一个受控信息包,并针对同一图像中可见的真实替代部件进行反事实重新询问。在三个OpenAI模型和1,068个预测中,提供目标区域而不提供其身份,使动作准确率保持在或低于0.53的多数基线(0.26、0.26和0.53),尽管模型在很大程度上复现了所提供的区域。仅提供身份而不提供位置则分别产生0.74、0.68和0.68的准确率。在此精选集中,所有高于基线的增益均出现在所提供的部件类别本身决定动作的情况下。无视觉对照使GPT-5的得分保持不变或提高,提供了与实质性类别到动作关联一致的证据。GPT-4o mini在一个条件下下降,因此这种解释并非普遍适用。添加关节类型和运动轴在六个模型层比较中并未提高准确率。在来自32个物体的74个反事实对中,GPT-5实现了0.86的成对加权合规率和0.07的捷径率,但在所有观察到的推升垂直案例中均失败。GroundBench识别了哪些提供的信息改变可供性行为,并测试看似有根据的性能是否可以通过文本捷径再现。
英文摘要
A companion evaluation found that naming the target part in a manipulation prompt increased action accuracy by 0.32-0.63 across eight vision-language models, with no model outperforming a constant baseline until the part was named. However, naming the part supplies information that a real system must infer, confounding visual grounding, mechanical reasoning, and category-to-action association. We introduce GroundBench, a diagnostic benchmark that separates these explanations through six branch-and-merge conditions, each adding a controlled information bundle, and a counterfactual re-ask targeting a real alternate part visible in the same image. Across three OpenAI models and 1,068 predictions, supplying the target region without its identity leaves action accuracy at or below the 0.53 majority baseline (0.26, 0.26, and 0.53), although the models largely reproduce the supplied region. Supplying identity without location instead yields 0.74, 0.68, and 0.68. Every above-baseline gain in this curated set occurs where the supplied part category itself determines the action. A no-vision control leaves GPT-5's scores unchanged or improved, providing evidence consistent with substantial category-to-action association. GPT-4o mini declines on one condition, so this interpretation is not universal. Adding joint type and motion axis does not improve accuracy across six model-stratum comparisons. On 74 counterfactual pairs from 32 objects, GPT-5 achieves 0.86 pair-weighted compliance with a 0.07 shortcut rate but fails all observed push-to-lift-vertical cases. GroundBench identifies which supplied information changes affordance behavior and tests whether apparently grounded performance can be reproduced through textual shortcuts.