RefineDrive:面向视觉-语言-动作驾驶的可靠失败引导学习
RefineDrive: Reliable Failure-Guided Learning for Vision-Language-Action Driving
浏览论文内容
中文总结 AI 辅助
提出RefineDrive,一种失败引导的后训练框架,通过可靠诊断、最小修正检索和安全分层强化学习,提升VLA驾驶模型的规划性能。
中文摘要 AI 辅助
用于自动驾驶的视觉-语言-动作(VLA)模型严重依赖成功的专家示范,导致模型特定的失败未被充分利用。从这些失败中学习受到不可靠的诊断、匹配不当的修正目标以及粗糙的奖励的阻碍。我们提出RefineDrive,一个失败引导的后训练框架,通过针对性的监督和安全感知的强化学习,从自生成的失败中学习。可靠诊断直接从模拟器状态中提取关于碰撞和可行驶区域违规的结构化、可验证的反馈。最小修正目标检索在聚类的轨迹库中搜索满足当前场景硬安全约束的邻近修正,优先保留失败预测的运动模式。以驾驶上下文和失败轨迹为条件,修正SFT学习生成诊断,随后将检索到的修正作为仅训练用的辅助任务。然后我们应用带有安全分层奖励的GRPO,该奖励严格优先考虑硬安全轨迹,对不安全轨迹和硬安全轨迹都保留连续的安全反馈,并且仅在硬安全满足后才奖励驾驶进展。在推理时,策略直接从驾驶上下文预测轨迹,无需显式的诊断或修复阶段。在NAVSIM v1上,RefineDrive将4B基础SFT策略的PDMS从87.7提升到91.7。使用相同的检查点且无需额外训练,RefineDrive在原始NAVTEST场景上使用NAVSIM v2扩展指标评估,达到89.4 EPDMS。受控消融实验支持结构化诊断监督、检索到的修正以及用于直接规划的安全分层优化的益处。
英文摘要
Vision-Language-Action (VLA) models for autonomous driving rely heavily on successful expert demonstrations, leaving model-specific failures underexploited. Learning from these failures is hindered by unreliable diagnoses, poorly matched correction targets, and coarse rewards. We propose RefineDrive, a failure-guided post-training framework that learns from self-generated failures through targeted supervision and safety-aware reinforcement learning. Reliable Diagnosis derives structured, verifiable feedback on collisions and drivable-area violations directly from simulator states. Minimum-Correction Target Retrieval searches a clustered human trajectory bank for nearby corrections that satisfy hard-safety constraints in the current scene, prioritizing preservation of the failed prediction's motion pattern. Conditioned on the driving context and failed trajectory, Correction SFT learns to generate the diagnosis followed by the retrieved correction as a training-only auxiliary task. We then apply GRPO with a Safety-Layered Reward that strictly prioritizes hard-safe trajectories, retains continuous safety feedback for both unsafe and hard-safe trajectories, and rewards driving progress only after hard safety is satisfied. At inference, the policy directly predicts trajectories from the driving context without an explicit diagnosis or repair stage. On NAVSIM v1, RefineDrive improves the 4B base SFT policy from 87.7 to 91.7 PDMS. Using the same checkpoint without additional training, RefineDrive achieves 89.4 EPDMS on the original NAVTEST scenes evaluated with NAVSIM v2 extended metrics. Controlled ablations support the benefits of structured diagnosis supervision, retrieved corrections, and safety-layered optimization for direct planning.
发表机构
- Zhejiang University(浙江大学)
- Yinwang Intelligent Technology Co., Ltd.(银望智能科技有限公司)
机构由 AI 辅助整理,请以论文原文为准。