如何在VLA后训练中使用(以及不使用)数据增强
How (and How Not) to Use Data Augmentation in VLA Post-Training
浏览论文内容
中文总结 AI 辅助
本研究系统探讨VLA后训练中的图像增强策略,发现仅增强评论家模块可显著提升分布外泛化性能,而增强演员会导致训练崩溃,并据此提出实用建议。
中文摘要 AI 辅助
视觉-语言-动作(VLA)模型目前在广泛的真实世界机器人任务中展现出强大的性能。然而,它们往往仍缺乏处理大型视觉分布外偏移的泛化能力。已有研究表明,通过强化学习(RL)对VLA进行后训练有助于提升鲁棒性,但仍有显著的改进空间。在这项工作中,我们系统地研究了图像增强对VLA后训练的影响。我们发现,在RL更新期间仅对评论家模块进行增强至关重要,而在 rollout 和更新期间保持演员的输入干净。对于π0.5和GR00T N1.5,这分别使LIBERO-Plus上的分布外成功率提高了7.8和10.0个百分点,而对演员进行增强则会导致训练完全崩溃。我们调查了一系列增强类型和强度,并为改善VLA后训练的泛化能力提供了实用建议。
英文摘要
Vision-language-action (VLA) models currently demonstrate strong performance in a wide range of real-world robotics tasks. However, they often still lack the generalization ability to handle large visual out-of-distribution shifts. Post-training of VLAs with reinforcement learning (RL) has been shown to benefit robustness, but significant room for improvement remains. In this work, we systematically study the effect of image augmentation on VLA post-training. We find that it is crucial to augment only the critic module during RL updates, while leaving the actor's input clean during both rollouts and updates. For $π_{0.5}$ and GR00T N1.5 this raises out-of-distribution success on LIBERO-Plus by $7.8$ and $10.0$ points respectively, while augmenting the actor collapses training entirely. We investigate a range of augmentation types and strengths, and provide practical recommendations for improving generalization in VLA post-training.
发表机构
- TU Eindhoven(埃因霍温理工大学)
机构由 AI 辅助整理,请以论文原文为准。