arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

CounterPlay:自我对弈驾驶策略的反事实后训练

CounterPlay: Counterfactual Post-Training for Self-Play Driving Policies

Jiarong Wei, Yin Wu, Runkai He, Abhinav Valada

arXiv 2609.21617首次发表:更新:

发表机构

University of Freiburg; CARIAD SE; Karlsruhe Institute of Technology(弗赖堡大学; CARIAD SE; 卡尔斯鲁厄理工学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

CounterPlay通过从失败任务回溯并在替代驾驶风格下重试,以反事实后训练提升自我对弈驾驶策略,在BehaviorBench上以1%训练预算取得最先进性能,并平衡任务完成与安全。

AI 中文摘要

在高吞吐量模拟器中进行自我对弈能够产生具有鲁棒闭环性能的驾驶策略,但随着训练规模的扩大,每单位模拟所带来的性能提升逐渐减少。策略在早期学会处理常见情况,而进一步的滚动模拟则反复遭遇未解决的失败。后训练为针对这些失败提供了机会,但现有方法主要是在已访问的状态下评估替代动作或后续延续,尽管成功的恢复可能需要在更早阶段改变驾驶风格。我们提出了CounterPlay,一种反事实的自我对弈后训练方法,它从失败的任务中回溯,并在替代驾驶风格下重试这些任务。CounterPlay基于三个关键组件。首先,失败驱动的回溯利用策略的价值估计来选择较早存储的状态,从该状态重试任务。其次,奖励条件化使得单一策略能够从该状态使用从谨慎到激进的候选风格重试任务。第三,CounterPlay仅在没有任何其他车辆相对于事实分支产生新的或更早的碰撞或偏离道路事件时,才保留完成任务的重试。通过新随机性验证的重试随后在部署条件下被蒸馏到策略中。在BehaviorBench上,CounterPlay在所有八种交通场景下的交互式和随机分割中均取得了最先进的分数,使用了10亿次后训练转换,这仅占锚点1000亿次自我对弈训练预算的1%。相对于锚点的改进在所有三种评估驾驶风格中均保持。CounterPlay解决了锚点在BehaviorBench上相当一部分超时案例,并实现了任务完成与安全性之间的平衡,这是持续自我对弈或采用更激进驾驶风格都无法达到的。

英文摘要

Self-play in high-throughput simulators yields driving policies with robust closed-loop performance, but improvement per unit of simulation diminishes as training scales. Policies learn to handle common situations early, while further rollouts repeatedly encounter unresolved failures. Post-training offers an opportunity to target these failures, but existing methods primarily evaluate alternative actions or continuations at visited states, although successful recovery may require changing driving style earlier. We propose CounterPlay, a counterfactual self-play post-training approach that backtracks from failed tasks and retries them under alternate driving styles. CounterPlay rests on three key components. First, failure-driven backtracking uses the policy's value estimates to select an earlier stored state from which to retry the task. Second, reward conditioning enables a single policy to retry the task from this state using candidate styles ranging from cautious to aggressive. Third, CounterPlay retains task-completing retries only if no other vehicle incurs a new or earlier collision or off-road event relative to the factual branch. Retries that pass verification with fresh randomness are then distilled into the policy under its deployment condition. On BehaviorBench, CounterPlay achieves state-of-the-art scores on both the Interactive and Random splits across all eight traffic regimes using 1B post-training transitions, which is just 1% of the anchor's 100B self-play training budget. Improvements over the anchor hold across all three evaluated driving styles. CounterPlay resolves a substantial fraction of the anchor's timeout cases on BehaviorBench and achieves a balance between task completion and safety that neither continued self-play nor adopting a more aggressive driving style attains.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑