arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

DriveReferee: 驾驶世界动作模型的几何安全判定无需学习

DriveReferee: Geometric Safety Verdicts Need Not Be Learned for Driving World-Action Models

Fengcheng Yu, Dhruv Parikh, Junjie Ye, Maulik Bhatt, Thang Vu, Igor Vasiljevic, Vitor Guizilini, Yue Wang

arXiv 2609.22762首次发表:更新:

发表机构

University of Southern California; Woven by Toyota; Toyota Research Institute(南加州大学; 丰田编织公司; 丰田研究院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

DriveReferee提出用学习的几何读出器预测场景状态,直接执行显式几何安全规则,替代学习验证器,在NAVSIM上以单摄像头输入达92.02 PDMS,无需外部训练数据。

AI 中文摘要

生成式世界动作模型(WAMs)联合生成未来视频和车辆动作,而其动作分支主要依靠专家模仿进行优化。然而,模仿并未为生成的轨迹提供明确的闭环几何判定,这使得在训练和部署期间的验证变得重要。闭环评估器可以检查碰撞和可行驶区域违规,但需要部署时不可用的特权场景状态。现有方法通常通过从传感器特征学习验证器来弥合这一差距。对于这些几何检查,规则本身是明确的。例如,碰撞由滚动自车足迹是否占据被占用的车辆空间决定。部署时不可用的是应用规则所需的场景状态。我们引入DriveReferee,它使用学习的几何读出器从相机观测预测场景表示,并直接执行几何安全规则而非学习它。由此产生的解析裁判从场景状态和候选轨迹评估碰撞和可行驶区域安全性。在训练期间,它基于真实状态对自采样轨迹评分,并将由此产生的偏好蒸馏到WAM策略中。在部署时,同一裁判基于预测状态评估生成的轨迹,并在需要时选择更安全的替代方案。解析裁判不需要针对判定的专门训练,其决策遵循明确的几何规则。在匹配的候选和推理预算下,它匹配或超越所有学习的验证器和启发式基线。给定相同的预测状态和轨迹,学习判定尽管需要数万个评估器标注的训练示例,但没有提供可衡量的下游收益。在完整的NAVSIM navtest上,DriveReferee在单摄像头视觉输入且无外部训练数据的情况下达到92.02 PDMS。

英文摘要

Generative world-action models (WAMs) jointly generate future video and vehicle actions, while their action branches remain primarily optimized by expert imitation. Yet imitation provides no explicit closed-loop geometric verdict for generated trajectories, making verification important during both training and deployment. Closed-loop evaluators can check collision and drivable-area violations, but require privileged scene state unavailable at deployment. Existing approaches often close this gap by learning a verifier from sensor features. For these geometric checks, the rule itself is explicit. For example, collision is determined by whether the rolled-out ego footprint overlaps occupied vehicle space. What is unavailable at deployment is the scene state needed to apply the rule. We introduce DriveReferee, which uses a learned geometry readout to predict the scene representation from camera observations and executes the geometric safety rule directly rather than learning it. The resulting analytic referee evaluates collision and drivable-area safety from a scene state and candidate trajectory. During training, it scores self-sampled trajectories on ground-truth state and distills the resulting preferences into the WAM policy. At deployment, the same referee evaluates generated trajectories on this predicted state and selects a safer alternative when needed. The analytic referee requires no verdict-specific training, and its decisions follow an explicit geometric rule. Under matched candidates and inference budgets, it matches or outperforms all learned-verifier and heuristic baselines. Given the same predicted state and trajectory, learning the verdict provides no measurable downstream gain despite requiring tens of thousands of evaluator-labeled training examples. On the full NAVSIM navtest, DriveReferee reaches 92.02 PDMS with single-camera visual input and no external training data.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑