arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

CCRV-Bench:基于约束的视觉语言模型因果推理评估

CCRV-Bench: Constraint-Based Evaluation of Causal Reasoning in Vision-Language Models

Linyuan Gao, Yuan Wu, Yi Chang

arXiv 2609.30979首次发表:更新:

发表机构

Jilin University(吉林大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出CCRV-Bench,一个基于约束的视觉语言模型因果推理评估基准,通过四个因果任务维度和多种约束减少捷径学习,实验表明约束鲁棒性与无约束性能无关。

AI 中文摘要

视觉语言模型(VLMs)在视觉任务中表现出色,但其视觉因果推理能力仍缺乏可靠的评估。现有评估难以区分模型是基于视觉证据进行因果推理,还是依赖统计相关性进行捷径学习,从而可能高估其实际能力。本文提出CCRV-Bench,一个针对单图像物理场景的约束驱动视觉因果推理基准。我们构建了一个正交框架,评估四个因果任务维度:因果关系发现、状态预测、因果诊断和干预。我们进一步引入实体符号化、空间定位、事实对抗约束和最小化输出约束,以减少捷径线索,同时保留任务所需的物理常识。对15个多模态模型的实验表明,约束敏感性因任务和模型而异:在四个因果任务中,干预的平均有效退化最大,空间定位是平均最具破坏性的约束,事实对抗约束提高了所有评估模型的DCR。这些结果表明,无约束性能并不能决定约束鲁棒性,单一的聚合分数可能掩盖因果识别、空间定位和符合约束表达方面的不同失败。CCRV-Bench为在受控约束下诊断基于图像的因果推理提供了标准化框架。代码可在以下网址获取:此https URL。

英文摘要

Vision-language models (VLMs) have demonstrated excellent performance in visual tasks, but their visual causal reasoning capabilities still lack reliable evaluation. Existing evaluations struggle to distinguish whether a model is performing causal reasoning based on visual evidence or relying on statistical correlations for shortcut learning, thereby potentially overestimating their actual capabilities. This paper proposes CCRV-Bench, a constraint-driven visual causal reasoning benchmark for single-image physical scenarios. We construct an orthogonal framework that evaluates four causal task dimensions: causal relation discovery, state prediction, causal diagnosis, and intervention-outcome prediction. We further introduce entity symbolization, spatial grounding, the factual adversarial constraint, and minimalist output constraints to reduce shortcut cues while preserving the physical commonsense required by the task. Experiments across 14 multimodal models show that constraint sensitivity is task- and model-dependent: intervention-outcome prediction has the largest average effective degradation among the four causal tasks, spatial grounding is the most damaging constraint on average, and the factual adversarial constraint improves DCR for all evaluated models. These results show that unconstrained performance does not determine constrained robustness and that a single aggregate score can obscure distinct failures in causal identification, spatial grounding, and constraint-compliant expression. CCRV-Bench provides a standardized framework for diagnosing image-grounded causal reasoning under controlled constraints. The code is available at https://github.com/0815linyuan/CCRVBench

Comments22 pages, 5 figures, 13 tables

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑