arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

自博弈驾驶中出现的现象与失效问题

What Emerges and What Breaks in Self-Play Driving

Laur Sisask, Ardi Tampuu, Tambet Matiisen

arXiv 2608.30819首次发表:更新:

发表机构

University of Tartu(塔尔图大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究在自博弈自动驾驶训练中,将模型扩展为Transformer,在CARLA和Waymax基准测试中表现不及Gigaflow,分析了失效模式、产生的交通规则及奖励条件对驾驶行为多样性的影响。

AI 中文摘要

近期,通过纯自博弈训练自动驾驶策略已取得良好成果。继Gigaflow和Puffer-Drive之后,我们以类似的自博弈方式训练驾驶策略,但将模型从多层感知机(MLPs)扩展为Transformer,并在真实城市的高清地图上进行训练,最终目标是将这些策略部署到该城市。在CARLA和Waymax基准测试中,我们的策略表现不及Gigaflow,我们将性能差距归因于特定的失效模式,包括在交通信号灯处的奖励作弊行为,以及缺乏在停车标志处停车的激励机制。我们进一步分析了自博弈中产生了哪些交通规则,以及这些规则与人类驾驶的匹配程度,还证实了奖励条件设置能产生预期的多样化驾驶行为。训练好的策略演示可在该https网址查看。

英文摘要

Training autonomous driving policies through pure self-play has recently shown promising results. Following Gigaflow and Puffer- Drive, we train driving policies in a similar self-play fashion, but extend the models from MLPs to Transformers and train on the high-definition map of a real city, where we ultimately aim to deploy them. On the CARLA and Waymax benchmarks, our policies fall short of Gigaflow, and we trace the gap to specific failure modes, including reward hacking at traffic lights and a missing incentive to stop at stop signs. We further analyze which traffic rules emerge from self-play and how closely they match human driving, and we confirm that reward conditioning yields the intended diversity of driving behaviors. A demonstration of a trained policy is available at https://laursisask-ut.github.io/eccvdemo.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑