arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.20747cs.ROcs.LG

MILER:非结构化自动驾驶中用于仿真到现实强化学习的语义中层表示

MILER: Semantic Mid-Level Representation for Sim-to-Real Reinforcement Learning in Unstructured Autonomous Driving

Thomas Steinecker, Denis Trescher, Alexander Bienemann, Thorsten Luettel, Mirko Maehlisch

首次发表
浏览论文内容

中文总结 AI 辅助

MILER通过语义中层表示模拟器和轨迹对齐策略,实现非结构化自动驾驶中感知与控制的零样本仿真到现实迁移,并在真实测试中验证了有效性。

中文摘要 AI 辅助

强化学习因其具有超人性能和自主学习策略的潜力,成为一种有前景的方法。然而,其在真实世界自动驾驶中的应用仍然很少,尤其是在非结构化环境中,这是因为非结构化环境的仿真到现实迁移面临挑战。在这项工作中,我们提出了MILER,一个具有零样本仿真到现实迁移能力的端到端策略框架。在离线训练期间,我们使用一个自定义的语义中层表示(MLR)模拟器,并通过强化学习训练策略网络,其控制输出直接应用于自行车模型。在真实车辆部署期间,相机和激光雷达数据由BEVFusion处理,以生成与MLR模拟器一致的语义鸟瞰图表示。策略网络生成的动作并不直接应用于真实车辆。相反,我们采用了一种轨迹对齐策略,使得感知和控制的零样本仿真到现实迁移成为可能。我们在一个包含众多挑战的多样化测试跑道上广泛评估了所提出的框架,这些挑战包括各种障碍物、发卡弯、高达33.6公里/小时的速度以及越野路段。总共,我们在一条3.0公里的测试跑道上使用两辆不同的车辆行驶了17.3公里,无需人工干预,从而证明了我们方法的有效性。此外,整个软件栈运行在Jetson AGX Orin上。

英文摘要

Reinforcement learning constitutes a promising approach owing to its potential for superhuman performance and self-learned policies. However, its application to real-world autonomous driving remains scarce, particularly in unstructured environments, because of the challenges associated with sim-to-real transfer for unstructured environments. In this work, we present MILER, an end-to-end policy framework with zero-shot sim-to-real transfer. During offline training, we employ a custom semantic mid-level representation (MLR) simulator and train the policy network using reinforcement learning, with its control outputs applied directly to a bicycle model. During deployment on the real vehicle, camera and LiDAR data are processed by BEVFusion to generate a semantic bird's-eye-view representation consistent with that of the MLR simulator. The actions generated by the policy network are not applied directly to the real vehicle. Instead, we employ a trajectory-alignment strategy that enables zero-shot sim-to-real transfer of both perception and control. We extensively evaluate the proposed framework on a diverse test track comprising numerous challenges, including various obstacles, hairpin curves, velocities of up to 33.6 km/h, and off-road sections. In total, we drove 17.3 km with two different vehicles on a 3.0 km test track without human intervention, thereby demonstrating the effectiveness of our approach. Furthermore, the entire software stack runs on a Jetson AGX Orin.

发表机构

  • University of the Bundeswehr Munich(慕尼黑联邦国防军大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑