看见并不足够:视觉语言模型感知证据但未能行动
Seeing is not Enough: Vision-Language Models Perceive Evidence but Fail to Act
- Florida State University(佛罗里达州立大学)
- Japan Advanced Institute of Science and Technology (JAIST)(日本先端科学技术大学院大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
视觉语言模型常感知到证据却决策错误,本文提出VPAC-Bench基准和SRT过程先验干预,证明边界对齐的过程先验能大幅降低模糊案例的过度承诺错误。
AI中文摘要:
视觉语言模型(VLM)在视觉问答基准测试中表现强劲,但常常做出与其已正确识别的视觉证据相矛盾的决策。我们将感知失败(相关证据未被识别)与过程失败(已识别的证据未能约束最终决策)区分开来。我们引入了VPAC-Bench,一个涵盖九个真实图像过程族(process families)的基准,每个图像都标注了其当前活动阶段及邻近的阶段转换。我们还提出了状态-相关性-目标(SRT),一系列结构化的过程先验干预措施,要求模型在回答前将可见证据与相关过程状态联系起来。在多个VLM中,过程失败普遍存在:正确列举视觉候选的模型在超过95%的模糊案例中仍过度承诺单一答案。一种显式的过程结构化干预将该比率降至13%以下,且不降低非模糊案例的性能。然而,过程先验的迁移是模型依赖的,通用SRT并未始终优于强链式思维(chain-of-thought)基线。当相关阶段转换已知时,边界对齐的SRT在组装、物理状态转换、导航与交通以及物体使用可供性任务中,显著优于通用过程提示和所有测试的链式思维基线。这些结果表明,当过程先验与场景的特定决策边界对齐时最为有用,这激励了面向过程基础视觉推理的边界感知先验选择。
英文摘要:
Vision-language models (VLMs) perform strongly on visual question answering benchmarks, yet often make decisions that contradict visual evidence they have already identified correctly. We distinguish perceptual failure, where relevant evidence is not recognized, from process failure, where recognized evidence fails to constrain the final decision. We introduce VPAC-Bench, a benchmark spanning nine real-image process families, with each image annotated by its current activity stage and nearby stage transition. We also propose State-Relevance-Target (SRT), a family of structured process-prior interventions that requires models to connect visible evidence to the relevant process state before answering. Across multiple VLMs, process failure is widespread: models that correctly enumerate visual candidates still over-commit to a single answer in more than 95% of ambiguous cases. An explicit process-structured intervention reduces this rate to below 13% without degrading performance on unambiguous cases. However, the transfer of process priors is model-dependent, and generic SRT does not consistently outperform strong chain-of-thought baselines. When the relevant stage transition is known, boundary-aligned SRT substantially outperforms generic process prompting and all tested chain-of-thought baselines across assembly, physical state transition, navigation and traffic, and object-use affordance tasks. These results show that process priors are most useful when aligned with the scene's specific decision boundary, motivating boundary-aware prior selection for process-grounded visual reasoning.