FIRE-VLA:面向自动驾驶的视觉-语言-动作模型的故障感知自演化
FIRE-VLA: Failure-Informed Self-Evolution for Vision-Language-Action Models in Autonomous Driving
浏览论文内容
中文总结 AI 辅助
该研究针对自动驾驶VLA模型,提出FIRE-VLA故障感知自演化框架,结合GRPO与自蒸馏,在nuScenes数据集上降低了规划误差与故障流行率,提升了模型性能。
中文摘要 AI 辅助
强化学习通过评估当前策略采样的轨迹来提升自动驾驶视觉-语言-动作(VLA)模型,群组相对策略优化(GRPO)利用每个回滚组内的奖励差异进行学习。当所有采样轨迹表现不佳时,该相对信号可对故障进行排名,但无法识别故障区域外的行为。我们提出FIRE-VLA,这是一种故障感知自演化框架,可将此类未解决的故障转化为下一策略的特权监督。低奖励、低多样性的组会触发对同一模型冻结的回合起始副本的自蒸馏,教师模型与学生模型参数规模相同,但仅教师模型可观测隐藏的未来轨迹。监督遵循学生生成的前缀,且仅限制于答案标记,同时GRPO仍在每个组中保持活跃。更新后的策略将作为下一回合的教师,使路由的故障分布随策略变化,无需更大的外部教师。从相同的Qwen2.5-VL-3B SFT检查点开始,比较匹配学生回滚和策略更新次数。在来自150个保留nuScenes场景的6019个示例上,FIRE-VLA保留了相当的单样本规划能力,将G=4平均L2从1.848米降至1.500米,并将评估持续故障流行率从13.03%降至11.20%。平均误差的降低主要源于罕见的严重回滚,而非普通轨迹的均匀改进。
英文摘要
Reinforcement learning improves autonomous-driving vision-language-action (VLA) models by evaluating trajectories sampled from the current policy. Group relative policy optimization (GRPO) learns from reward differences within each rollout group. When all sampled trajectories are poor, this relative signal can rank failures without identifying behavior outside the failed region. We introduce FIRE-VLA, a failure-informed self-evolution framework that converts such unresolved failures into privileged supervision for the next policy. Low-reward, low-diversity groups trigger self-distillation from a frozen round-start copy of the same model. Teacher and student have the same parameter scale, but only the teacher observes the hidden future trajectory. Supervision follows the student's generated prefix and is restricted to answer tokens, while GRPO remains active for every group. The updated policy supplies the teacher for the next round, allowing the routed failure distribution to change with the policy without requiring a larger external teacher. Starting from the same Qwen2.5-VL-3B SFT checkpoint, the comparison matches student rollout and policy-update counts. On 6,019 examples from 150 held-out nuScenes scenes, FIRE-VLA retains comparable single-sample planning, reduces G=4 mean L2 from 1.848 to 1.500 m, and lowers evaluation-persistent failure prevalence from 13.03% to 11.20%. The reduction in mean error arises mainly from rare severe rollouts rather than uniform improvement across ordinary trajectories.
发表机构
- Harbin Institute of Technology(哈尔滨工业大学)
机构由 AI 辅助整理,请以论文原文为准。