arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

他们看到了什么?通过视觉-语言-动作模型的视角解读复杂道路场景以实现安全可靠的自动驾驶学习

What Do They See? Interpreting Complex Road Scenarios Through the Eyes of Vision-Language-Action Models for Safe and Trustworthy Autonomous Vehicle Learning

Kalpana Panda, Wesley Maia, Vinti Agarwal, Ross Greer

arXiv 2607.16938首次发表:更新:

发表机构

Birla Institute of Technology and Science, Pilani; University of California, Merced(贝拉理工科学学院皮拉尼分校; 加州大学默塞德分校)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究自动驾驶模型内部逻辑不透明问题,提出反事实消融框架CVAA,通过移除摄像头图像中物体创建反事实集评估模型响应差异,应用于阿尔帕马约1轨迹预测器,分析物体因果影响,还通过可解释性技术深入理解模型,创建可解释驾驶系统巩固人机信任。

AI 中文摘要

端到端自动驾驶模型如今能够在复杂道路场景中导航,将原始传感器观测直接映射到观测路径以进行开环评估,在闭环评估中也往往能有效驾驶。然而,由于交通场景的复杂性,这些安全关键系统的内部逻辑在很大程度上仍不透明。我们提出了一个名为反事实视觉动作分析(CVAA)的反事实消融框架,利用逼真的生成式修复技术系统地从前置摄像头图像中移除单个检测到的物体,以准备反事实集来评估模型响应的差异。这分离了每个物体存在对模型规划行为的因果效应。应用于跨越210个nuScenes驾驶场景的阿尔帕马约1轨迹预测器,我们创建了一个数据集Counter - nuScenes,从中发现模型“路径”内的车辆和行人如预期主导因果影响,而交通灯相对于其图像占比产生了不成比例的影响。但我们也发现模型对人类驾驶员认为无关的物体有强烈反应的情况。这引发了一个更深层次的问题:模型是将场景视为影响结果的单个物体的总和,还是编码了一组与人类可读场景元素不对应的完全不同的内部特征?为进一步理解这一点,我们使用机制可解释性技术比较原始图像对和修复后图像对的中间表示,并检查通过各个模型层移除的效果。这两个阶段共同提供了一条从行为审计到表示理解的路径,创建可解释的驾驶系统并巩固人类与人工智能之间的信任。

英文摘要

End-to-end autonomous driving models are now able to navigate complex road scenarios, mapping raw sensor observations directly to observed paths for open-loop evaluation and often effective driving in closed-loop evaluation. Yet the internal logic of these safety-critical systems remains largely opaque, due to the complexity of traffic scenes. We propose a counterfactual ablation framework called Counterfactual Vision Action Analysis (CVAA) that systematically removes individual detected objects from front-camera images using photorealistic generative inpainting to prepare counterfactual sets to evaluate the difference in the model's response. This isolates the causal effect of each object's presence on the model's planning behaviour. Applied to the Alpamayo 1 trajectory predictor across 210 nuScenes driving scenes, we create a dataset Counter -nuScenes, using which we see that vehicles and pedestrians within the model's 'path' dominate causal influence as expected, while traffic lights, as expected, exert disproportionate effect relative to their image footprint. However, we also find cases where the model responds strongly to objects a human driver would consider irrelevant. This brings forth a deeper question: does the model itself view the scene as a sum of individual objects influencing the outcome, or does it encode an entirely different set of internal features that do not correspond to human-legible scene elements? To further understand this, we compare intermediate representations of original and inpainted image pairs using mechanistic interpretability techniques and examine the effect of the removal through the various model layers. Together, these two stages offer a path from behavioral auditing to representational understanding, creating explainable driving systems and solidifying human-AI trust.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑