AI 中文总结
研究在大规模视觉-语言-动作模型中,通过将离线监督纳入强化学习,结合离线与在线训练优点,提升训练效率。经实验验证,该混合方法在保持强大分布外能力的同时,所需训练预算减半,实现效率与性能双赢。
AI 中文摘要
通常观察到,在线强化学习(RL)在广泛的性能指标上比离线方法产生性能更好的策略。特别是,RL训练的策略表现出更强的分布外(OOD)行为,而仅用模仿学习方法训练的模型往往难以做到。最近一项研究引入了一个以OOD为重点的基准,并报告说RL训练的视觉-语言-动作(VLA)策略比通过监督微调(SFT)训练的对应策略具有明显更好的OOD性能和略好的分布内(IND)性能。在这项工作中,我们研究混合离线-在线训练是否能结合两种方法的优点。具体而言,我们研究通过离线数据或离线训练的参考策略进行离线监督正则化的RL方法。我们在OOD基准上评估这些方法,并将它们与仅离线训练和标准RL进行比较。我们的结果表明,虽然离线训练本身实现的OOD性能有限,但将离线监督纳入RL可保持强大的OOD能力,同时大幅提高训练效率。特别是,引导方法达到了接近标准RL的性能,同时所需训练预算大约减半。混合方法在实现这种效率提升的同时,并没有在速度和OOD性能之间进行权衡,而是保留了强大的OOD能力。项目页面:this https URL
英文摘要
It is commonly observed that online reinforcement learning (RL) produces better-performing strategies than offline methods across a broad range of performance measures. In particular, RL-trained policies exhibit stronger out-of-distribution (OOD) behavior, where models trained only with imitation learning approaches often struggle. A recent study introduced an OOD-focused benchmark and reported that RL-trained vision-language-action (VLA) policies achieve noticeably better OOD performance and slightly better in-distribution (IND) performance than their counterparts trained with supervised fine-tuning (SFT). In this work, we investigate whether hybrid offline-online training can combine the advantages of both approaches. Specifically, we study RL methods regularized by offline supervision via either offline data or an offline-trained reference policy. We evaluate these approaches on the OOD benchmark and compare them with both offline-only training and standard RL. Our results show that although offline training achieves limited OOD performance by itself, incorporating offline supervision into RL preserves strong OOD capability while substantially improving training efficiency. In particular, the guided methods reach performance close to that of standard RL while requiring roughly half of the training budget. Rather than producing a trade-off between speed and OOD performance, the hybrid approach retains strong OOD capability while achieving this efficiency gain. Project page: https://alstar8.github.io/offline-supervision-vla-rl