arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.36438cs.RO

World4Scorer:面向自动驾驶的基于结果的世界建模

World4Scorer: Outcome-Grounded World Modeling for Autonomous Driving

Jieyuan Pei, Meiyi Lu, Sining Ang, Yubo Zhao, Zhangyi Hu, Mingwei Xu, Haokai Ding, Wei Li, Zihan You, Jianwei Zheng, Li Yu, Yifeng Pan, Ji Tao, Rongjunchen Zhang, Yan Wang

首次发表
浏览论文内容

中文总结 AI 辅助

World4Scorer提出基于模拟器结果监督的轨迹条件JEPA预测器,为自动驾驶生成-选择规划器评分未执行计划,在NAVISIM-v2和闭环Bench2Drive上达到最先进性能。

中文摘要 AI 辅助

自动驾驶需要在周围交通演变过程中选择安全且高效的计划。生成-选择规划器提出多条轨迹并对其进行评分以供执行,这类规划器在NAVSIM基准上已优于具有代表性的直接预测基线。其评分器必须比较从未执行过的计划。驾驶日志仅记录已执行轨迹的未来,因此匹配日志中的未来可能使对备选方案的预测不受约束;相比之下,模拟器可以为每个候选方案标注结果。我们提出World4Scorer,将评分器构建为轨迹条件下的JEPA风格预测器:它为每个候选方案预测一个状态,并从中读取该候选方案的评分。模拟器结果标签监督所有候选方案的状态,而已执行轨迹的观测未来将预测器锚定到真实场景演变。由于一个预测器生成每个候选方案的状态,该锚定可以约束用于对未执行计划评分的共享参数,而未来本身仅在训练时需要。生成的候选方案大多评分良好,因此场景匹配的候选库将低评分计划纳入结果监督;逐帧选择可能冲突,因此惯性重排序保持连续选择的一致性。World4Scorer在NAVISIM-v2上实现了最先进的性能,并在闭环Bench2Drive上取得了强大的适配系统结果。在固定LeWM世界模型和规划预算的情况下,基于结果的评分还改善了OGBench-Cube基准上的操作规划。

英文摘要

Autonomous driving requires choosing a safe and efficient plan as surrounding traffic evolves. Generate-and-select planners propose multiple trajectories and score them for execution, and they have outperformed representative direct-prediction baselines on NAVSIM. Their scorer must compare plans that were never executed. Driving logs record the future of only the executed trajectory, so matching the logged future can leave predictions for the alternatives unconstrained; a simulator, in contrast, can label the outcome of every candidate. We introduce World4Scorer, which builds the scorer as a trajectory-conditioned JEPA-style predictor: it predicts a state for each candidate and reads the candidate's scores from that state. Simulator outcome labels supervise the states of all candidates, and the observed future of the executed trajectory anchors the predictor to real scene evolution. Because one predictor produces every candidate's state, the anchor can constrain shared parameters used to score unexecuted plans, while the future itself is needed only during training. Generated candidates mostly score well, so a scene-matched bank adds low-scoring plans to the outcome supervision; framewise choices can conflict, so inertial re-ranking keeps consecutive selections consistent. World4Scorer achieves state-of-the-art NAVSIM-v2 performance and a strong adapted-system result on closed-loop Bench2Drive. With the LeWM world model and planning budget fixed, outcome-based scoring also improves manipulation planning on the OGBench-Cube benchmark.

发表机构

  • Institute for AI Industry Research (AIR), Tsinghua University(清华大学人工智能产业研究院(AIR))
  • HiThink Research(海天瑞声研究)
  • The Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州))
  • Zhejiang University(浙江大学)
  • University of Science and Technology of China(中国科学技术大学)
  • SMBU(深圳北理莫斯科大学)
  • University of Washington(华盛顿大学)
  • Mohamed bin Zayed University of Artificial Intelligence(穆罕默德·本·扎耶德人工智能大学)
  • Zhejiang University of Technology(浙江工业大学)
  • Southeast University(东南大学)
  • Changan Automobile(长安汽车)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑