发表机构
KE:SAI; ETH Zürich; NVIDIA Research; Stanford University; ELLIS Institute Tübingen(KE:SAI; 苏黎世联邦理工学院; 英伟达研究院; 斯坦福大学; ELLIS图宾根研究所)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
OPTED提出一种无渲染教师策略的在线微调方法,通过特权教师提供监督,在仿真中高效提升端到端驾驶性能,显著减少交互次数。
AI 中文摘要
随着仅扩展预训练数据带来的收益递减,后训练在自动驾驶等物理AI领域中变得越来越重要。端到端驾驶策略通过行为克隆在人类演示上进行开环预训练。然而,闭环部署过程中的复合误差可能使车辆偏离训练数据分布,增加安全关键事件的风险。闭环后训练可以缓解这一风险,但对于基于传感器的策略需要昂贵的仿真。我们提出OPTED(端到端驾驶的在线策略微调),它将强化学习与端到端策略的后训练解耦:一个特权教师使用RL在向量化输入(高清地图和边界框)上训练。然后,该教师在闭环后训练期间为预训练的学生提供监督。我们将OPTED应用于两个基于摄像头的模型TransFuser和VaVAM,并在AlpaSim中使用真实驾驶日志的神经重建(3DGS)进行微调。驾驶得分分别提高了1.6倍和9.5倍。在受控实验中,OPTED以比直接RL后训练少约三个数量级的仿真交互次数达到相同的闭环性能,同时更接近人类先验。项目页面:此https URL
英文摘要
As scaling pre-training data alone yields diminishing returns, post-training is becoming increasingly important across physical AI domains such as autonomous driving. End-to-end driving policies are pre-trained in open loop with behavior cloning on human demonstrations. However, compounding errors during closed-loop deployment can take the vehicle outside the training data distribution, increasing the risk of safety-critical incidents. Closed-loop post-training can mitigate this risk but requires costly simulation for sensor-based policies. We propose OPTED (on-policy fine-tuning for end-to-end driving) which decouples reinforcement learning from the post-training of the end-to-end policy: a privileged teacher is trained using RL on vectorized inputs (HD-map and bounding boxes). This teacher then provides supervision to the pre-trained student during closed-loop post-training. We apply OPTED to two camera-based models, TransFuser and VaVAM, and fine-tune them in AlpaSim, using neural reconstructions (3DGS) of real driving logs. Driving scores increase by factors of 1.6$\times$ and 9.5$\times$, respectively. In controlled experiments OPTED matches closed-loop performance with approximately three orders of magnitude fewer simulator interactions than direct RL post-training, while staying closer to the human prior. Project page: https://01dami23.github.io/opted/
Comments9 pages, 5 figures