arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

TerraZero:用于大规模零演示自博弈的过程式驾驶模拟

TerraZero: Procedural Driving Simulation for Zero-Demonstration Self-Play at Scale

Zhouchonghao Wu, Akshay Rangesh, Weixin Li, Wei-Jer Chang, Zachary Lee, Saeed Bonab, Tim Wang, Wei Zhan

arXiv 2607.13028首次发表:更新:

发表机构

Applied Intuition(应用直觉)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究旨在训练强大自动驾驶智能体,提出TerraZero过程式驾驶模拟器和自博弈训练堆栈,其速度快、逼真且多样。通过特定方式填充地图生成无限场景,智能体仅靠强化学习训练,能零样本泛化,在多个基准测试中表现出色。

AI 中文摘要

训练强大的自动驾驶智能体需要一个足够快以进行大规模强化学习、足够逼真以基于真实世界地图结构确定行为、足够多样以覆盖日志数据很少包含的安全关键长尾情况的模拟器。我们提出了TerraZero,一种过程式驾驶模拟器和自博弈训练堆栈。一个可配置的C引擎在CPU上运行模拟,在GPU上通过零拷贝路径进行策略推理,在单个服务器级GPU上每秒维持130万个智能体步,比现有对象级模拟器快得多,同时保持保真度。TerraZero仅将日志数据作为真实世界地图几何的来源,用随机的基于规则的道路使用者和信号控制器填充每个地图,并在每集随机化智能体动力学、奖励和大小,因此一张地图会产生无限的场景集。每个报告的策略仅通过跨GPU的计算高效自博弈方法进行强化学习从零开始训练,推理时无需人工演示和后备规划器。策略在城市和数据集之间进行零样本泛化,包括在没有明确监督的情况下出现的左侧交通驾驶。作为自我策略,TerraZero是第一个在InterPlan长尾基准测试中名列前茅的完全学习策略,优于更大规模的学习规划器;在常规驾驶val14上,它是最佳方法之一且最安全,具有最佳的碰撞和碰撞时间分数。在Waymo Open Sim Agents逼真度方面,相同方法优于其他无演示方法,并与最强的基于参考的自博弈方法竞争。一个堆栈同时服务于两个角色:跨汽车和卡车动力学的驾驶策略,以及联合控制车辆、行人和骑自行车者的模拟智能体。

英文摘要

Training robust autonomous driving agents requires a simulator fast enough for reinforcement learning at scale, realistic enough to ground behavior in real-world map structure, and diverse enough to cover the safety-critical long tail that logged data rarely contains. We present TerraZero, a procedural driving simulator and self-play training stack that meets these goals. A configurable C engine runs simulation on the CPU and policy inference on the GPU over a zero-copy path, sustaining 1.3M agent-steps per second on a single server-grade GPU, far faster than existing object-level simulators, while keeping fidelity lighter single-agent systems omit: heterogeneous agents, multiple dynamics models, and full traffic-rule enforcement. TerraZero uses logged data only as a source of real-world map geometry, populating each map with randomized rule-based road users and signal controllers and randomizing agent dynamics, rewards, and sizes per episode, so one map yields an effectively unbounded set of scenarios. Every reported policy trains from scratch by reinforcement learning alone, with zero human demonstrations, no imitation, no logged trajectories, and no fallback planner at inference, on a compute-efficient self-play recipe scaled across GPUs. The policies generalize zero-shot across cities and datasets, including emergent left-hand-traffic driving without explicit supervision. As an ego policy, a single checkpoint is, to our knowledge, the first fully learned policy to top both val14 and the interactive long-tail InterPlan suite. On Waymo Open Sim Agents realism the same recipe outperforms other demonstration-free methods and is competitive with the strongest reference-anchored self-play method. One stack serves both roles: state-of-the-art demonstration-free driving policies across dynamics for cars and trucks, and sim agents that jointly control vehicles, pedestrians, and cyclists.

CommentsTechnical Report from Applied Intuition Research

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑