多种成功路径:面向VLA泛化的多样性驱动RL微调
Many Ways to Succeed: Diversity-Driven RL Fine-Tuning for VLA Generalization
- Nanjing University(南京大学)
- Harbin Institute of Technology (Shenzhen)(哈尔滨工业大学(深圳))
- Australian National University(澳大利亚国立大学)
- University of Technology Sydney(悉尼科技大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
提出DRIVE方法,将成功行为多样性作为显式RL目标,通过分组比较轨迹并给予内部奖励,提升VLA策略的域外泛化,在多个基准和真实平台上显著提升OOD成功率。
AI中文摘要:
强化学习(RL)微调通过闭环经验提升视觉-语言-动作(VLA)策略,但其在微调分布之外的泛化能力仍然有限。我们的分析揭示了探索的选择性重塑:RL在全局上收缩行为,却使成功轨迹多样化,以更少的rollout获得成功,并且比监督微调覆盖了更多潜在任务有效解空间。更广泛的成功模式覆盖可能在分布偏移下提供替代策略。受此启发,我们提出DRIVE(面向VLA泛化的多样性驱动RL微调),将成功行为的多样性转化为显式的RL目标。DRIVE在匹配的任务条件下对rollout进行分组,通过时间对齐比较其轨迹,并从相对行为多样性中推导出成功条件的内部奖励。该设计鼓励更广泛地覆盖可行解,而不奖励多样的失败或表面的时间差异。在LIBERO-Plus、ManiSkill3和RoboTwin 2.0上,DRIVE将平均域外(OOD)性能相比普通RL微调提升了5.3个百分点(在$\pi_0$上)和2.0个百分点(在$\pi_{0.5}$上)。在双臂AgileX PiPER-X平台上,DRIVE进一步将平均OOD成功率从64.1%提升至73.3%(+9.2个百分点),展示了在物理部署下持续存在的收益。
英文摘要:
Reinforcement learning (RL) fine-tuning improves vision-language-action (VLA) policies through closed-loop experience, yet generalization beyond the fine-tuning distribution remains limited. Our analysis reveals a selective reshaping of exploration: RL contracts behavior globally, yet diversifies successful trajectories, elicits success with fewer rollouts, and covers more of the latent task-valid solution space than supervised fine-tuning. Broader successful-mode coverage may provide alternative strategies under distribution shifts. Inspired by this, we introduce DRIVE (Diversity-driven RL fIne-tuning for VLA gEneralization), which turns successful-behavior diversity into an explicit RL objective. DRIVE groups rollouts under matched task conditions, compares their trajectories with temporal alignment, and derives a success-conditioned intrinsic reward from relative behavioral diversity. This design encourages broader coverage of feasible solutions without rewarding diverse failures or superficial timing differences. Across LIBERO-Plus, ManiSkill3, and RoboTwin 2.0, DRIVE improves the average out-of-domain (OOD) performance over vanilla RL fine-tuning by 5.3 points on $π_0$ and 2.0 points on $π_{0.5}$. On a dual-arm AgileX PiPER-X platform, DRIVE further increases average OOD success from 64.1% to 73.3% (+9.2 points), demonstrating gains that persist under physical deployment.