EDCT-Bench:通过解释驱动的反事实测试揭示视觉语言模型中的忠实性差距
EDCT-Bench: Uncovering Faithfulness Gaps in VLMs via Explanation-Driven Counterfactual Testing
- Mercedes-Benz Research & Development North America(梅赛德斯-奔驰北美研发中心)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
提出EDCT协议与EDCT-Bench基准,通过反事实测试揭示VLM解释与视觉证据的忠实性差距,并证明其生成数据对微调有高价值。
AI中文摘要:
视觉语言模型(VLMs)能够生成听起来合理但与它们所引用的视觉证据不一致的自然语言解释(NLEs)。我们提出了解释驱动的反事实测试(EDCT),这是一种基于干预的协议,它提取模型解释中引用的视觉概念,对其应用经过验证的最小编辑,并测试由此产生的答案和解释是否与编辑后的图像保持一致。利用这一协议,我们创建了EDCT-Bench,一个涵盖三个互补领域的全面基准:知识密集型视觉问答(OK-VQA)、安全关键的驾驶(DriveLM)和3D空间推理(3DSRBench)。在评估的VLMs中,EDCT揭示了显著的忠实性差距,模型经常产生与经过验证的视觉变化不一致的响应。最后,我们的微调研究表明,EDCT生成的反事实提供了高影响力的训练信号。
英文摘要:
Vision-Language Models (VLMs) can produce Natural Language Explanations (NLEs) that sound plausible yet remain inconsistent with the visual evidence they cite. We present Explanation-Driven Counterfactual Testing (EDCT), an intervention-based protocol that extracts visual concepts cited in a model's explanation, applies verified minimal edits to them, and tests whether the resulting answer and explanation remain consistent with the edited image. Using this protocol, we create EDCT-Bench, a comprehensive benchmark spanning three complementary domains: knowledge-intensive visual question answering (OK-VQA), safety-critical driving (DriveLM), and 3D spatial reasoning (3DSRBench). Across the evaluated VLMs, EDCT reveals substantial faithfulness gaps, with models frequently producing responses inconsistent with verified visual changes. Finally, our fine-tuning study suggests that EDCT-generated counterfactuals provide high-impact training signals.