arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

RoG-DAgger:面向端到端驾驶的基于回滚的后训练方法

RoG-DAgger: Rollout-Guided Post-Training for End-to-End Driving

Liangyu Zhong, Joachim Sicking, Fabian Hueger, Hanno Gottschalk

arXiv 2608.24525首次发表:更新:

发表机构

CARIAD SE, Volkswagen Group; Institute of Mathematics, Technical University of Berlin(大众集团旗下CARIAD SE; 柏林工业大学数学研究所)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究针对端到端驾驶系统训练与推理不匹配的问题,提出 RoG-DAgger 后训练框架,通过回滚优化专家演示与监督,在多个驾驶基准上显著提升了模型的驾驶性能与成功率。

AI 中文摘要

近期端到端驾驶系统在闭环基准测试中展现出较强性能,但仍主要通过开环模仿学习在固定的专家采集数据上进行训练。这种训练与推理的不匹配导致策略在策略诱导状态下易受影响,累积的误差可能引发安全关键型故障。克服该问题的一种有前景的后训练方法是数据集聚合(DAgger),其会在策略诱导状态下收集专家演示,随后在聚合后的数据集上对策略进行微调。然而,现有的驾驶 DAgger 流程面临三大挑战:i)专家受限于有限的轨迹与速度解空间;ii)相对于即将到来的故障,接管可能发生得过早或过晚;iii)特权专家决策可能依赖于学生无法获取的信息。为解决这些问题,我们提出 RoG-DAgger,一种使用短程运动学回滚在安全关键状态下构建高质量专家演示的后训练框架。具体而言,RoG-DAgger 扩展了专家的轨迹与速度解空间,并通过回滚评估候选计划以构建预防性监督;此外,它利用回滚可解性在估计的无法挽回点附近确定接管时机;最后,它使专家的视野与学生的视野对齐,以提供与学生兼容的监督。在分布内(包括长程)和分布外评估中,RoG-DAgger 提升了端到端模型 SimLingo 的性能:在 Bench2Drive 上驾驶得分提高 5.3 分、成功率提高 6.2 个百分点,在 Longest6 v2 上驾驶得分从 22 翻倍至 44,在 Fail2Drive 上的分布外成功率从 55% 提升至 66%。

英文摘要

Recent end-to-end driving systems demonstrate strong performance on closed-loop benchmarks, yet are still predominantly trained on fixed expert-collected data using open-loop imitation learning. This training-inference mismatch leaves the policy vulnerable in policy-induced states, where accumulated errors can lead to safety-critical failures. A promising post-training approach to overcome this issue is Dataset Aggregation (DAgger), which gathers expert demonstrations in policy-induced states and subsequently fine-tunes the policy on the resulting aggregated dataset. Existing driving DAgger pipelines, however, face three challenges: i) the expert is restricted to a limited trajectory-and-speed solution space, ii) takeover may occur too early or too late relative to impending failures, and iii) privileged expert decisions may rely on information unavailable to the student. To address this, we introduce RoG-DAgger, a post-training framework that uses short-horizon kinematic rollouts to construct high-quality expert demonstrations in safety-critical states. Specifically, RoG-DAgger expands the expert's trajectory-and-speed solution space and evaluates candidate plans through rollout to construct preventive supervision. Moreover, it uses rollout solvability to time the takeover near the estimated point of no return. Lastly, it aligns the expert's field of view with that of the student to provide student-compatible supervision. Across in-distribution (including long-horizon) and out-of-distribution evaluations, RoG-DAgger improves the end-to-end model SimLingo by 5.3 driving-score points and 6.2 percentage points in success rate on Bench2Drive, doubles its driving score from 22 to 44 on Longest6 v2, and improves out-of-distribution success rate from 55\% to 66\% on Fail2Drive.

Commentspreprint, under review

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑