arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

有时你必须先跑才能走:用于VLM自动驾驶的Run-then-Walk调度策略

Sometimes You Gotta Run Before You Can Walk: Run-then-Walk Scheduling Strategy for VLM Autonomous Driving

Yuqi Ye, Shangkun Sun, Junhong Lin, Jiayi Zhao, Changhao Peng, Wei Zheng, Guoqing Liu, Tiesong Zhao, Wei Gao

arXiv 2609.25831首次发表:更新:

发表机构

Peking University; Minieye Technology Co., Ltd; Fuzhou University(北京大学; 深圳佑驾创新科技股份有限公司; 福州大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对VLM自动驾驶中GRPO强化学习效率与安全难以兼顾的问题,提出Run-then-Walk两阶段奖励调度策略,先探索进度再修复安全,在多个基准上提升性能并减少40%-50%训练轮次。

AI 中文摘要

近期基于VLM的自动驾驶规划器采用GRPO风格的强化学习来优化驾驶性能。然而,现有的GRPO方案要么优化驾驶效率,冒着追求进度但不安全行为的风险,要么强制执行早期安全约束,导致过于保守的行为;两者都需要长时间的训练。为了解决这些问题,我们首先揭示了两种不同的RL机制:一种积极探索高进度的进度机制(Run-GRPO),以及一种在稳定进度下恢复安全的安全机制(Walk-GRPO)。基于这一发现,我们提出了Run-then-Walk,一种简单而有效的两阶段奖励调度策略,用于GRPO,实现了更好的性能和更快的收敛。与单阶段RL(可能在单个训练阶段内专注于进度、安全或两者的混合)不同,这种调度明确地将进度发现与安全修复分开。在Run阶段,我们专注于进度,使策略能够摆脱保守偏差并发现高进度模式。在随后的Walk阶段,我们引入端点和安全策略来修复Run阶段的不安全行为。这种反向调度克服了Walk优先方法的保守性和联合优化的不安全进度追求。我们使用多种基于VLM的规划器在多个基准上进行了验证:NAVSIMv1、NAVSIMv2、Navhard和nuScenes。大量实验表明,驾驶性能得到提升,同时所需的RL训练轮次比基线少40%-50%。

英文摘要

Recent VLM-based autonomous driving planners adopt GRPO-style reinforcement learning to optimize driving performance. However, existing GRPO recipes either optimize driving efficiency, risking progress-seeking but unsafe behavior, or enforce early safety constraints, leading to overly conservative behavior; both require lengthy training. To solve these problems, we first reveal two distinct RL regimes: a progress regime (Run-GRPO) that aggressively explores high progress, and a safety regime (Walk-GRPO) that restores safety under stable progress. Based on this finding, we propose $\textit{Run-then-Walk}$, a simple yet effective two-stage reward scheduling strategy for GRPO, achieving both better performance and faster convergence. Unlike one-stage RL, which may focus on progress, safety, or a mixture of both within a single training phase, this schedule explicitly separates progress discovery from safety repair. In the $\textit{Run}$ phase, we focus on progress, allowing the policy to escape the conservative bias and discover high-progress modes. In the subsequent $\textit{Walk}$ phase, we introduce endpoint and safety strategy to repair unsafe behaviors from the Run phase. This reversed schedule overcomes the conservatism of Walk-first methods and the unsafe progress-seeking of joint optimization. We validate it with various VLM-based planners on multiple benchmarks: NAVSIMv1, NAVSIMv2, Navhard, and nuScenes. Extensive experiments demonstrate improved driving performance while requiring 40--50\% fewer RL training epochs than the baselines. Code is available at https://github.com/haha-yuki-haha/AutoDrive-P3_with_Run-then-walk.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑