arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.22582cs.CVcs.AI

超越排行榜:端到端与VLA驾驶策略在域偏移下的反事实诊断

Beyond the Leaderboard: Counterfactual Diagnosis of End-to-End and VLA Driving Policies Under Domain Shift

Ruolin Yang, Zilin Huang, Buoyue Wang, Zhengyang Wan, Yuhao Luo, Zihao Sheng, Sikai Chen

首次发表
浏览论文内容

中文总结 AI 辅助

针对端到端和VLA驾驶策略,提出基于反事实帧编辑的诊断方法,揭示排行榜排名无法预测新场景行为,并评估策略的安全性、响应性和稳定性。

中文摘要 AI 辅助

端到端和视觉-语言-动作(VLA)驾驶策略通过排行榜排名进行比较,但排名只报告结果,而非背后的行为,因此它无法很好地预测策略在新地点将如何表现。在六个已发布的策略上,基于nuScenes开环误差或NAVSIM排行榜的排名并不能推广到新地点中行人靠近自车走廊的场景。我们提出了一种反事实检查方法:几百个真实帧,每个帧以两种方式编辑(移除行人,或施加夜间风格的扰动),每次编辑都由独立的检测器验证,并将规划轨迹的变化解读为诊断而非分数。从这些编辑中读取两个因果轴,并基于它们构建五项考试,以区分分数所合并的内容:策略计划行驶多远,看到行人是否带来安全性,该响应是否随危险程度缩放,当无需响应时计划是否移动,以及无关的照明变化使其移动多少。在246个NAVSIM近行人场景中,在行人位于规划路径上的单元内,仅1.9%的响应是真正的避让,而在我们的开环协议下,每个策略的中位净空变化至多为0.03米,规划距离的中位变化至多为0.08米。在一项从左侧驾驶到右侧驾驶的预注册测试中,暴露度和特异性排序、照明判定及碰撞结果得以迁移,而点值和危险敏感性判定则未迁移。作为选择报告解读,这些档案表明哪个策略因规划较短而安全,哪个策略覆盖类似人类的距离但不让行,哪个策略在无需响应的变化下不稳定,并为每个判定定价:大多数在几十帧内稳定,危险敏感性需要数百帧。代码和编辑后的帧将发布。

英文摘要

End-to-end and vision-language-action (VLA) driving policies are compared by leaderboard rank, but a rank reports an outcome, not the behaviour behind it, so it predicts poorly how a policy will behave at a new site. On six released policies, rank on nuScenes open-loop error or on NAVSIM's leaderboard does not carry over to scenes with a pedestrian near the ego corridor at a new site. We propose a counterfactual check-up: a few hundred real frames, each edited two ways (pedestrian removed, or re-lit by a night-style perturbation), every edit verified by an independent detector, and the change in the planned trajectory read as a diagnosis rather than a score. From these edits two causal axes are read, and five exams built on them separate what a score merges: how far the policy plans to drive, whether seeing the pedestrian buys safety, whether that response scales with danger, whether the plan moves when nothing requires it, and how much an irrelevant lighting change moves it. On 246 NAVSIM near-pedestrian scenes, in the cells where the pedestrian lies on the planned path only 1.9% of responses are genuine avoidance, and under our open-loop protocol the median clearance change is at most 0.03 m and the median change in planned distance at most 0.08 m for every policy. In a pre-registered test from left- to right-hand drive, the exposure and specificity orderings, the lighting verdict and the collision outcome transfer, while point values and the hazard-sensitivity verdict do not. Read as a selection report, the profiles say which policy is safe because it plans short, which covers a human-like distance without yielding, and which is unsteady under a change that requires no reaction, and they price each verdict: most settle within a few dozen frames, hazard sensitivity needs hundreds. Code and edited frames will be released.

发表机构

  • University of Wisconsin–Madison(威斯康星大学麦迪逊分校)

机构由 AI 辅助整理,请以论文原文为准。

↑