AI 中文总结
本文针对零样本视觉语言控制中VLMs决策是否基于视觉输入的问题,通过多组消融实验分析了多种VLMs的表现,提出对称共识守护者方法,验证了VLMs可作为有界的危险助手。
AI 中文摘要
视觉语言模型(VLMs)越来越多地被用作零样本控制器,但成功的轨迹并不一定表明决策是基于视觉输入的:模拟器动态和保守的动作先验可以在没有有意义感知的情况下产生有利的分数。我们通过一组输入消融实验对此进行研究:盲图像控制、重复相同输入、车道轴反射、非视觉基线以及流水线完整性检查。我们分析了9种直接动作模型、6种结构化局部VLMs以及探索性VLM-MPC层次结构,在2种 embodiment 和3种模拟器上的32874次评分调用。直接控制的结果大多是负面的:恒定慢速策略优于程序化几何控制器,多个模型是图像不变或几乎恒定的,识别纵向危险的模型在反射下仍然无法转换LEFT和RIGHT。没有局部VLM满足纵向和横向grounding的联合标准。然而,仅图像的确定性正控制以0.090米的MAE估计前车距离,且具有精确的镜像等变性,确认刺激携带足够的视觉信息;失败是模块化的,而非普遍的。事后的、泄漏控制的对称共识守护者从16个校准帧中选择两个模型,并在原始和反射视图中冻结4选2的危险投票。在272个保留帧上,它达到0.954的平衡准确率(剧集聚类自助法95%置信区间[0.895,0.990]);嵌套留一剧集法在所有12折中都恢复了相同的对和阈值。在平局时弃权(不执行)将承诺的平衡准确率提高到0.973,覆盖率为0.824。在确定性感知保留横向权限的情况下,离线模块化重播达到0.934的动作一致性和精确的镜像等变性。这些结果支持当前VLMs作为有界、选择性的危险助手,而非整体的零样本控制器。
英文摘要
Vision-language models (VLMs) are increasingly used as zero-shot controllers, but successful trajectories do not necessarily show that decisions are grounded in visual input: simulator dynamics and conservative action priors can produce favourable scores without meaningful perception. We investigate this with an input-ablation battery: blind-image controls, repeated identical inputs, lane-axis reflection, non-visual baselines, and pipeline-integrity checks. Across nine direct-action models, six structured local VLMs, and an exploratory VLM-MPC hierarchy, we analyse 32,874 scored calls over two embodiments and three simulators. The direct-control results are largely negative: a constant-SLOW policy outperforms a scripted geometric controller, several models are image-invariant or nearly constant, and models that recognize longitudinal hazards still fail to transform LEFT and RIGHT under reflection. No local VLM meets the joint longitudinal and lateral grounding criteria. However, an image-only deterministic positive control estimates the lead gap with 0.090 m MAE and exact mirror equivariance, confirming the stimuli carry sufficient visual information; the failures are modular, not universal. A post-hoc, leakage-controlled symmetry-consensus guardian selects two models from 16 calibration frames and freezes a 2-of-4 hazard vote across original and reflected views. On 272 held-out frames it reaches 0.954 balanced accuracy (episode-cluster bootstrap 95% CI [0.895,0.990]); nested leave-one-episode-out recovers the same pair and threshold in all 12 folds. Abstaining on ties raises committed balanced accuracy to 0.973 at 0.824 coverage. With deterministic perception retaining lateral authority, offline modular replay achieves 0.934 action agreement and exact mirror equivariance. These results support current VLMs as bounded, selective hazard assistants, not monolithic zero-shot controllers.