arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.32157cs.CVcs.RO

CausalDriveBench:评估自动驾驶视觉-语言-行动模型中的因果推理

CausalDriveBench: Evaluating Causal Reasoning in Vision-Language-Action Models for Autonomous Driving

Narendiran Chembu, Navvrat Rao, Shreedhar Shreeshail Kodate, Gayatri Srujana Banda, Arko Sarkar, Abhinav Khanna, Rajarshee Das, Umesh Kanala, Siddarth Khandelwa… 展开作者

Narendiran Chembu, Navvrat Rao, Shreedhar Shreeshail Kodate, Gayatri Srujana Banda, Arko Sarkar, Abhinav Khanna, Rajarshee Das, Umesh Kanala, Siddarth Khandelwal, Kumar Aman, Aish Dubey, Kaustubh Beedkar, Arjun Jain

首次发表
浏览论文内容

中文总结 AI 辅助

提出CausalDriveBench基准,基于Pearl因果层级评估自动驾驶VLA模型的因果推理,发现最佳模型QA准确率仅70.6%,且因果QA与轨迹准确率不相关,表明流畅理由和准确轨迹不证明因果理解。

中文摘要 AI 辅助

用于自动驾驶的视觉-语言-行动(VLA)模型在生成预测轨迹的同时产生自然语言推理,但这种推理是否反映场景的因果结构仍未得到验证。我们引入了CausalDriveBench,一个基于Pearl因果层级(PCH)的评估框架,通过结构化视觉问答(QA)和替代轨迹预测来测试驾驶特定VLA中的因果推理。为此,我们在nuScenes上构建了因果场景图,区分因果活跃、休眠和干扰实体,将感知显著性从因果相关性中分离出来。该基准涵盖PCH的所有四个层级(关联、干预和反事实以及因果发现)用于QA生成。对于较高层级,我们还在指定场景修改下提供参考轨迹,实现补充推理层面评估的行动层面验证。总体而言,该基准包含从nuScenes导出的7,285个经过验证的因果QA对和1,000个反事实轨迹。我们评估了10个驾驶特定VLA和3个通用VLM,并报告了三个发现。首先,最佳模型仅达到70.6%的QA准确率,13个模型中有4个得分低于随机概率。其次,将每个驾驶VLA与共享其语言骨干的通用VLM进行比较,驾驶微调的成本在因果QA上从2到34个百分点不等,后训练设计解释了这一差异。第三,因果QA和轨迹准确率在模型间统计上不相关:在反事实提示下,预测轨迹要么过度反应,要么坍缩到观测场景基线。综合来看,这些结果表明,流畅的理由或准确的观测场景轨迹都不构成因果理解的证据。

英文摘要

Vision-Language-Action (VLA) models for autonomous driving produce natural-language reasoning alongside predicted trajectories, but whether this reasoning reflects the causal structure of the scene remains untested. We introduce CausalDriveBench, an evaluation framework grounded in Pearl's Causal Hierarchy (PCH) that tests causal reasoning in driving-specific VLAs through structured visual question answering (QA) and alternative-trajectory prediction. To this end, we construct causal scene graphs over nuScenes that distinguish causally active, dormant, and distractor entities, separating perceptual salience from causal relevance. The benchmark spans all four rungs of PCH (association, intervention, and counterfactual along with causal discovery) for QA generation. For the higher rungs, we additionally provide reference trajectories under specified scene modifications, enabling action-level verification that complements reasoning-level evaluation. In total, the benchmark contains 7,285 verified causal QA pairs and 1,000 counterfactual trajectories derived from nuScenes. We evaluate 10 driving-specific VLAs and 3 general-purpose VLMs, and report three findings. First, the best model reaches only 70.6% QA accuracy, and 4 of 13 models score below random chance. Second, comparing each driving VLA to the general-purpose VLM that shares its language backbone, the cost of driving fine-tuning ranges from 2 to 34 percentage points on causal QA, with post-training design explaining the spread. Third, causal QA and trajectory accuracy are statistically uncorrelated across models: under counterfactual prompts, predicted trajectories either over-react or collapse onto the observed-scene baseline. Taken together, these results show that neither fluent rationales nor accurate observed-scene trajectories constitute evidence of causal understanding.

↑