arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

面向多风格端到端驾驶的长时一致且感知交互的世界模型

Long-Horizon Consistent and Interaction-Aware World Models for Multi-Style End-to-End Driving

Yuxuan Han, Kunyuan Wu, Liyunong Yang, Zilu Wang, Cansen Jiang, Yi Xiao, Liang Hu

arXiv 2609.03225首次发表:更新:

发表机构

School of Intelligence Science and Engineering, Harbin Institute of Technology, Shenzhen; Autonomous Driving Center, Shanghai Utopilot Technology Co.Ltd.(哈尔滨工业大学(深圳)智能科学与工程学院; 上海元戎智能科技有限公司自动驾驶中心)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对现有世界模型在自动驾驶中长时一致性、交互建模及风格适应性的局限,提出StyleDrive框架,通过三项改进提升性能,在Bench2Drive基准上获显著提升并实现仿真到真实的迁移。

AI 中文摘要

端到端自动驾驶越来越多地采用基于世界模型的强化学习框架,通过“想象回滚”提升学习效率。然而现有世界模型存在三大核心局限:长时想象回滚中的时间不一致性、自车-环境交互建模不足、对多样化驾驶风格的适应性有限。为应对这些挑战,我们提出StyleDrive,一种基于世界模型的学习框架,在统一学习范式中联合强化长时一致性、显式解耦交互交通状态,并支持多风格策略优化。首先,我们引入时间一致性正则化,通过门控交叉注意力整合历史隐状态,稳定长时想象回滚并缓解误差累积。其次,设计显式状态解耦模块,将与自车相关的交互状态与无关状态分离,在复杂交通场景中实现更具可解释性和高效的决策。第三,通过Group Relative Policy Optimization实现多风格驾驶行为,该方法用轨迹级相对优势替代每步奖励优化,降低奖励方差且无需重新训练即可支持多样化驾驶风格。我们在Bench2Drive闭环驾驶基准上评估StyleDrive,取得88.44的驾驶得分(较此前最优基于世界模型的方法提升17.08)和66.82的成功率(提升16.58)。此外,我们将StyleDrive部署在真实自动导引车平台上,在动态驾驶场景中展现出良好的仿真到真实的迁移能力。

英文摘要

End-to-end autonomous driving has increasingly adopted world model-based reinforcement learning frameworks to improve learning efficiency through \textit{imagined rollouts}. However, existing world models suffer from three key limitations: temporal inconsistency in long-horizon imagined rollouts, inadequate modeling of ego-environment interactions, and limited adaptability to diverse driving styles. To address these challenges, we propose \textit{StyleDrive}, a world-model-based learning framework that jointly enforces long-horizon consistency, explicitly disentangles interactive traffic states, and supports multi-style policy optimization within a unified learning paradigm. First, we introduce a temporal consistency regularization that integrates historical latent states through gated cross-attention, stabilizing long-horizon imagined rollouts and mitigating error accumulation. Second, we design an explicit state disentanglement module that separates ego-relevant from ego-irrelevant interactive states, enabling more interpretable and efficient decision-making in complex traffic scenarios. Third, we enable multi-style driving behaviors through Group Relative Policy Optimization, which replaces per-step reward optimization with trajectory-wise relative advantages, reducing reward variance and supporting diverse driving styles without retraining. We evaluate StyleDrive on the Bench2Drive closed-loop driving benchmark, achieving a driving score of 88.44 (+17.08 over the previous best world model-based method) and a success rate of 66.82 (+16.58). Furthermore, we deploy StyleDrive on a real automated guided vehicle platform and demonstrate promising sim-to-real transfer capability in dynamic driving scenarios.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑