arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.38862cs.ROcs.AI

高效多模态规划:面向自动驾驶的奖励引导偏好优化

Efficient Multi-Modal Planning with Reward-Guided Preference Optimization for Autonomous Driving

Chenglin Chen, Lujia Wang, Xinhu Zheng, Jun Ma, Haoang Li

首次发表
浏览论文内容

中文总结 AI 辅助

提出EMPlan,一种结合稀疏锚点与偏移细化模块的高效多模态轨迹规划方法,通过奖励引导微调提升安全性,在NAVSIM基准上实现精度与效率的平衡。

中文摘要 AI 辅助

安全高效的轨迹规划在自动驾驶中至关重要。然而,现有的端到端方法在计算效率和安全性保障方面往往存在不足。基于模仿学习的方法易受因果混淆影响,而基于规则的评分方法通常带来沉重的计算开销,并面临目标不一致的问题。此外,基于偏好的方法依赖严格的成对标注,限制了数据利用率。为克服这些局限,我们提出了EMPlan,一种由奖励引导微调驱动的高效多模态轨迹规划方法。我们设计了一种混合架构,将稀疏锚点与偏移细化模块相结合,以实现高效的多模态轨迹预测。稀疏锚点以低延迟提供粗略的轨迹候选,随后由偏移模块进行细化,以提高预测精度。为在不增加推理成本的情况下增强安全性,我们采用了两阶段训练范式,包括预训练和奖励引导微调。在微调阶段,我们利用基于规则的奖励信号和非成对偏好监督,将预训练策略优化为更安全的轨迹选择。我们在非反应式NAVSIM基准上评估了EMPlan,它在规划精度和效率之间取得了良好平衡,在实时约束下展现了优越性能。

英文摘要

Safe and efficient trajectory planning is essential in autonomous driving. However, existing end-to-end approaches often fall short in both computational efficiency and safety guarantees. Methods based on imitation learning suffer from causal confusion, while rule-based scoring approaches often incur heavy computational overhead and suffer from objective misalignment. Additionally, preference-based methods rely on strict pairwise annotations, limiting data utilization. To overcome these limitations, we propose EMPlan, an efficient multi-modal trajectory planning method powered by reward-guided fine-tuning. We design a hybrid architecture that combines sparse anchors with an offset refinement module for efficient multi-modal trajectory prediction. Sparse anchors provide coarse trajectory candidates with low latency, which are subsequently refined by the offset module for higher prediction accuracy. To enhance safety without incurring additional inference costs, we adopt a two-stage training paradigm consisting of pretraining and reward-guided fine-tuning. During fine-tuning, we leverage rule-based reward signals and unpaired preference supervision to refine the pretrained policy toward safer trajectory selection. We evaluate EMPlan on the non-reactive NAVSIM benchmark, where it strikes a favorable balance between planning accuracy and efficiency, demonstrating superior performance under real-time constraints.

发表机构

  • The Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州))

机构由 AI 辅助整理,请以论文原文为准。

↑