arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.37046cs.CV

盲点中的速度:面向自动驾驶的VLMs动态感知可解释性分析

Speed in the Blind Spot: An Interpretability Analysis of Dynamic Perception in VLMs for Autonomous Driving

Katharina Winter, Stefan Englmeier, Fabian B. Flohr

首次发表
浏览论文内容

中文总结 AI 辅助

本研究通过速度理解诊断任务,分析自动驾驶中视觉-语言模型的动态感知能力,发现其内部表示与口头输出存在差距,且驾驶专业化虽增强运动表示,但不保证可靠的时间基础。

中文摘要 AI 辅助

视觉-语言模型越来越多地应用于自动驾驶系统,但其从视觉输入中恢复动态物理状态的能力仍未得到充分表征。我们将速度理解作为一项受控诊断,在三个任务上进行研究:周围智能体速度、当前自车速度以及短期未来自车速度提议。在nuScenes数据集上,我们评估了开放权重的通用型和PhysicalAI视觉-语言模型,以及面向驾驶的Alpamayo-1.5视觉-语言-动作模型,使用了多种输入和输出公式。我们将口头评估与时间扰动、反事实自车速度提示和隐藏表示的线性探针相结合。这些任务表现出不同的失败模式。周围智能体速度以智能体特定的形式被弱编码,而当前自车速度通常在内部可访问但口头表达不佳:连续探针达到4.7-5.8 km/h的平均绝对误差(MAE),而口头输出的平均绝对误差为10.2-16.8 km/h。多帧输入相对于单帧输入在口头表达上的提升不一致,且帧顺序很少被利用。在非优化的QLoRA下,任务特定适应改善了任务相关的潜在速度表示和口头读出,但连续周围智能体速度估计仍然较弱,而大多数未来速度增益在帧洗牌后仍然存在,表明时间基础有限。驾驶专用的Alpamayo-1.5在周围智能体和未来自车速度的潜在表示上表现更强,而当前自车速度的可解码性相当,且探针与口头表达之间的差距仍然显著。因此,驾驶专业化可以加强运动表示,但并不能保证在场景和自车状态上都有更强的编码或可靠的读出。结果表明,合理的规划输出并不一定意味着对底层动态状态的可靠恢复或时间基础。

英文摘要

Vision-Language Models are increasingly used in autonomous-driving systems, yet their ability to recover dynamic physical state from visual input remains insufficiently characterized. We study velocity understanding as a controlled diagnostic across three tasks: surrounding-agent speed, current ego speed, and short-horizon future ego-speed proposal. On nuScenes, we evaluate open-weight general-purpose and PhysicalAI VLMs, together with the driving-oriented Alpamayo-1.5 Vision-Language-Action model, using multiple input and output formulations. We combine verbal evaluation with temporal perturbations, counterfactual ego-speed hints and linear probes of hidden representations. The tasks exhibit distinct failure modes. Surrounding-agent speed is weakly encoded in an agent-specific form, whereas current ego speed is often internally accessible but poorly verbalized: continuous probes achieve 4.7-5.8 km/h MAE compared with 10.2-16.8 km/h MAE for verbal outputs. Multiple frames provide inconsistent verbal gains to single frame inputs, and frame order is rarely exploited. Under non-optimized QLoRA, task-specific adaptation improves both task-relevant latent speed representations and verbal readout, but continuous surrounding-agent speed estimation remains weak, while most future-speed gains survive frame shuffling, indicating limited temporal grounding. Driving specialized Alpamayo-1.5 shows stronger latent representations for surrounding-agent and future ego speed, while current ego-speed decodability is comparable and substantial probe-verbal gaps remain. Thus, driving specialization can strengthen motion representations but does not guarantee stronger encoding across both scene and ego states or reliable readout. The results show that plausible planning outputs do not necessarily imply reliable recovery or temporal grounding of the underlying dynamic state.

发表机构

  • Munich University of Applied Sciences(慕尼黑应用科学大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑