AI 中文总结
RoboFFT提出一种前向过程强化学习框架,通过加噪和加权分数/流匹配损失微调生成式机器人策略,在模拟和真实任务中提升性能、稳定性和训练效率。
AI 中文摘要
生成模型,如扩散模型和基于流的模型,通过从示范中捕获复杂和多模态的动作分布,在机器人策略学习方面显示出强大的潜力。然而,仅通过模仿学习训练的策略常常受到不完美示范和分布偏移的影响,而进一步的改进通常需要额外的专家数据。强化学习通过环境交互提供了一种自然的解决方案,但由于似然估计的难解性,有效微调生成式机器人策略仍然具有挑战性。在这项工作中,我们提出了RoboFFT,一种用于微调生成式机器人策略的前向过程强化学习框架,该框架对采样的动作应用前向加噪,并使用加权的分数/流匹配损失来构建用于PPO风格更新的代理策略比率。我们在代表性的模拟基准上评估了RoboFFT与流行的生成式机器人策略,包括长视距规划和稀疏奖励设置。大量的实验和分析表明,RoboFFT在实现更好的稳定性和训练效率的同时,持续提高了性能。我们进一步将RoboFFT集成到一个真实世界的强化学习框架中,并展示了其在真实世界任务中的有效性。项目网站:此https URL。
英文摘要
Generative models, such as diffusion and flow-based models, have shown strong promise for robot policy learning by capturing complex and multimodal action distributions from demonstrations. However, policies trained solely with imitation learning often suffer from imperfect demonstrations and distributional shifts, while further improvement typically requires additional expert data. Reinforcement learning offers a natural solution through environment interaction, but effectively finetuning generative robot policies remains challenging due to the intractability of likelihood estimation. In this work, we propose RoboFFT, a forward-process reinforcement learning framework for finetuning generative robot policies, which applies forward noising to sampled actions and uses the weighted score / flow matching loss to construct a surrogate policy ratio for PPO-style updates. We evaluate RoboFFT with popular generative robot policies on representative simulation benchmarks, including long-horizon planning and sparse reward settings. Extensive experiments and analysis demonstrate that RoboFFT consistently improves performance while achieving better stability and training efficiency. We further integrate RoboFFT into a real world RL framework and demonstrate its effectiveness in real world tasks. Project website: https://student-of-holmes.github.io/RoboFFT/.
CommentsAccepted by CoRL2026. Project Website: https://student-of-holmes.github.io/RoboFFT/