arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.30122cs.CVcs.AIcs.LG

面向视觉-语言-动作(VLA)驾驶的多轨迹监督与策略优化对齐

Aligning Multi-Trajectory Supervision with Policy Optimization for VLA Driving

Tian Zhang, Zhuo Huang, Hongrui Ye, Yu Wu, Zengmao Wang, Kaixuan Zhou

首次发表
浏览论文内容

中文总结 AI 辅助

针对VLA驾驶中多轨迹与GRPO不对齐的问题,本文提出对齐框架,引入两种互补机制,在NAVSIM数据集上取得优于GRPO基线的性能。

中文摘要 AI 辅助

视觉-语言-动作(VLA)驾驶方法日益将多轨迹模仿学习与组相对策略优化(GRPO)相结合,使得轨迹选择对最终性能至关重要。然而,一些能提升模仿效果的高分轨迹会通过诱导与当前策略可行行为分布不对齐的优势估计,降低后续GRPO的效果,导致策略更新偏离安全合规的行为。为解决该问题,本文提出一种将多轨迹监督与策略优化对齐的新型框架。针对可行区域外不可行噪声轨迹引发的策略梯度偏差,本文将增强轨迹约束在真实可行区域的邻近流形内,采用帕累托最优准则替代传统聚合分数,仅保留非支配候选,从源头过滤冲突样本。为确保扩展轨迹监督在策略优化中被有效吸收,本文引入两种互补机制:可行性优先优势分配与动态蒸馏。前者将帕累托信用适配至每个回滚组的可行性构成,引导完全不可行组趋向安全参考;后者在多轮优化中更新教师轨迹,持续传递有用监督。两者共同将扩展监督的益处逐步转化为策略提升。在NAVSIM v1和v2数据集上,本文方法在单轨迹推理下分别达到91.4 PDMS和89.1 EPDMS,且在658个初始失败场景中恢复440个,比原始GRPO基线高11.1%。

英文摘要

Vision-language-action (VLA) driving methods increasingly combine multi-trajectory imitation learning with group-relative policy optimization (GRPO), making trajectory selection critical to final performance. However, some high-scoring trajectories that improve imitation can degrade subsequent GRPO by inducing advantage estimates misaligned with the current policy's feasible behavior distribution, driving updates away from safe and compliant behaviors. To address this, we propose a novel framework that aligns multi-trajectory supervision with policy optimization. To address the policy gradient bias induced by infeasible noisy trajectories outside the feasible region, augmented trajectories are constrained to a neighboring manifold of the ground-truth feasible region, and a Pareto-optimality criterion is adopted in place of the conventional aggregate score, retaining only non-dominated candidates and thereby filtering out conflicting samples at the source. To ensure that expanded trajectory supervision is effectively absorbed during policy optimization, we introduce two complementary mechanisms: feasibility-first advantage assignment and dynamic distillation. The former adapts Pareto credit to the feasibility composition of each rollout group and guides fully infeasible groups toward safe references. The latter updates teacher trajectories across refinement rounds to continually transfer useful supervision. Together, they progressively translate the benefits of expanded supervision into policy improvement. On NAVSIM v1 and v2, our method achieves 91.4 PDMS and 89.1 EPDMS, respectively, under single-trajectory inference, and recovers 440 of 658 initially failed scenes, 11.1\% higher than the original GRPO baseline.

发表机构

  • School of Computer Science, Wuhan University(武汉大学计算机学院)
  • Dongfeng Research & Development Institute(东风研发院)

机构由 AI 辅助整理,请以论文原文为准。

↑