发表机构
The University of Texas at Austin; Bosch Center for AI(德克萨斯大学奥斯汀分校; 博世人工智能中心)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究提出CL4AD,将课程学习集成到批量自动驾驶模拟器中,在GPUDRIVE实验中使自动驾驶RL智能体的训练效率大幅提升,提前10亿步达到99%成功率,减少77%的时间,样本效率提升67%。
AI 中文摘要
近期,自动驾驶的批量模拟器已支持大规模训练强化学习(RL)智能体,在数天内即可涵盖数千种交通场景及数十亿次交互。尽管这类高吞吐量数据馈送使RL算法的训练速度远超以往,但样本效率并未同步提升:标准训练方案采用域随机化均匀采样场景,导致大量交互消耗于对学习贡献极小的案例。课程学习可通过自适应优先选择对策略改进最关键的场景来解决该问题。我们提出CL4AD,首次将课程学习与批量自动驾驶模拟器集成,将场景选择建模为无监督环境设计问题。我们引入效用函数,结合智能体的成功率、行为真实性及现有遗憾估计函数来构建课程。在GPUDRIVE上开展的大规模实验表明,课程学习比域随机化提前10亿步达到99%的成功率,将 wall-clock 时间减少77%,且在除最大规模外的所有情况下均优于基于静态和动态属性的启发式课程。在计算资源有限的消融实验中,课程学习使样本效率提升67%。我们还研究了效用函数在大规模场景下的表现,以及训练过程中优先选择的场景如何演变。我们在GPUDRIVE中发布了CLForAD的实现代码。
英文摘要
Batched simulators for autonomous driving have recently enabled training reinforcement learning (RL) agents at scale, encompassing thousands of traffic scenarios and billions of interactions within a matter of days. Although such high-throughput feeds RL algorithms faster than ever, their sample-efficiency has not kept pace: As the standard training scheme, domain randomization uniformly samples scenarios, thereby consuming a vast number of interactions on cases that contribute little to learning. Curriculum learning offers a remedy by adaptively prioritizing scenarios that matter most to policy improvement. We present CL4AD, the first integration of curriculum learning into batched autonomous driving simulators by framing scenario selection as an unsupervised environment design problem. We introduce utility functions that shape curricula based on success rates and the realism of the agent's behavior, in addition to existing regret-estimation functions. Large-scale experiments in GPUDRIVE demonstrate that curriculum learning achieves a 99% success rate a billion steps earlier than domain randomization, reducing wall-clock time by 77%, and outperforms heuristic curricula with static and dynamic attributes, with only one exception at the largest scale. An ablation under limited compute shows that curriculum learning improves sample efficiency by 67%. We also investigate how utility functions behave at scale, and how prioritized scenarios evolve during training. We release an implementation of CLForAD in GPUDRIVE.
Comments31 pages, 18 figures. Under review at NeurIPS 2026