arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.13410cs.RO

用于零样本跨底盘自适应自动驾驶的自我动力学增强世界模型

Ego-Dynamics-Augmented World Model for Autonomous Driving with Zero-Shot Cross-Embodiment Adaptation

Zhidong Wang, Jingsong Liang, Zirui Li, Zhan Chen, Han Yu, Chen Lv

首次发表
浏览论文内容

中文总结 AI 辅助

研究针对自动驾驶中基于世界模型的强化学习问题,提出DynaDreamer方法,通过增强自我动力学先验改进世界模型,减少自我运动建模负担,实现零样本跨底盘自适应,实验证明该方法显著提升驾驶任务成功率。

中文摘要 AI 辅助

基于世界模型(WM)的强化学习通过在潜在空间中想象长距离轨迹,实现了样本高效的端到端自动驾驶学习。然而,大多数驾驶WM在本质上以自我为中心的鸟瞰图(BEV)表示上运行,导致自我运动与场景动力学纠缠。本文提出DynaDreamer,一种动力学增强的Dreamer风格强化学习方法,通过明确的自我动力学先验增强WM来解决此问题。物理信息自我动力学编码器 - 解码器提取自我状态历史到紧凑且可识别的上下文,调节因果Transformer WM的先验和后验潜在。在想象过程中,自我动力学预测器向前传播此上下文以保持自我动力学先验与展开同步。信息理论分析表明,基于此上下文进行条件设定可减少观察转换的预测熵和先验 - 后验Kullback - Leibler散度,将WM的建模负担限制在自我运动之外的场景动力学上。此外还具有零样本跨底盘自适应能力。实验表明,DynaDreamer在城市和高速公路驾驶场景中分别比最强基线提高任务成功率28%和61%,在推断到未见底盘时优势升至73%。

英文摘要

End-to-end autonomous driving requires generalization ability across platforms with dissimilar physical characteristics. The chassis defines the physical embodiment of each platform, and real-world fleets span sub-tonne microcars to bus-class vehicles. Consequently, the driving stack must either be retrained per platform or adapt to the underlying chassis dynamics online. World model (WM)-based reinforcement learning offers a sample-efficient path toward end-to-end autonomous driving on egocentric bird's-eye-view (BEV) representations, but its effectiveness hinges on how faithfully the WM captures the ego vehicle's dynamics. This work identifies a structural bottleneck in BEV-based WMs: observation transitions entangle ego-motion with scene dynamics, consuming modeling capacity at the cost of imagination accuracy. This burden is embodiment-dependent: dissimilar chassis produce different observation warps under the same control input. The proposed DynaDreamer addresses this bottleneck by conditioning the WM's latent distributions on a physics-informed ego-dynamics context derived from a lateral dynamics model with a neural tire force formulation. This context is extracted online via a neural-ODE encoder-decoder that simultaneously identifies the underlying chassis parameters. Information-theoretic analysis confirms that this conditioning removes the ego-motion terms from both the WM's transition entropy and its prior-posterior KL divergence. The identified physical parameterization enables zero-shot cross-embodiment adaptation across a dynamically diverse fleet without per-platform retraining. Simulation results show 28% and 43% improvements in driving task success rates over the strongest baseline in urban and highway scenarios, and the advantage over the base Transformer WM reaches up to 73% when extrapolating to unseen chassis.

发表机构

  • School of Mechanical and Aerospace Engineering, Nanyang Technological University(南洋理工大学机械与宇航工程学院)
  • Collaborative Initiative, Interdisciplinary Graduate Programme, Nanyang Technological University(南洋理工大学跨学科研究生项目合作计划)
  • College of Computing and Data Science, Nanyang Technological University(南洋理工大学计算与数据科学学院)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑