发表机构
Joy Future Academy, JD; School of Artificial Intelligence, Jilin University(京东探索研究院; 吉林大学人工智能学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
HaWMPO提出幻觉感知世界模型策略优化方法,通过估计生成图像可靠性并抑制不可靠动作,提升VLA策略在长期任务中的成功率,在LIBERO和真实机器人上取得显著提升。
AI 中文摘要
通用机器人策略在机器人操作任务中展现出强大的泛化能力,但在复杂的长期任务场景中,其成功率仍然有限。近期方法通过在线强化学习在真实机器人上改进视觉-语言-动作(VLA)策略,但此类训练依赖昂贵的物理交互,样本效率低,并可能带来硬件和安全风险。世界模型通过支持想象回放中的策略优化,提供了一种有前景的替代方案。然而,世界模型生成的长期回放常常遭受预测幻觉问题,产生有偏的状态转移,从而误导策略学习。为解决这一问题,我们提出了幻觉感知世界模型策略优化(HaWMPO),一种用于VLA策略与世界模型后训练的闭环强化学习流程。具体而言,HaWMPO引入了一种动作条件下的幻觉感知模型,用于估计生成图像序列的可靠性,并通过奖励软机制将幻觉分数纳入组相对策略优化中,在训练期间抑制不可靠的动作块。在LIBERO基准上,HaWMPO取得了最佳平均成功率,相比基础模型提升15.0%,相比最强基线提升2.8%;在G1机器人上的真实世界实验进一步验证了其有效性,将两项操作任务的平均成功率从67.5%提升至80.0%。
英文摘要
Generalist robot policies have demonstrated strong generalization across robotic manipulation tasks, yet their success rates remain limited in com- plex long-horizon scenarios. Recent methods improve Visual-Language-Action (VLA) policies through online reinforcement learning on real robots, but such training relies on costly physical interactions, suffers from low sample efficiency, and may introduce hardware and safety risks. World models offer a promising alternative by enabling policy optimization with imagined rollouts. However, long-horizon rollouts generated by world models often suffer from prediction hal- lucinations, producing biased state transitions that can mislead policy learning. To address this issue, we propose Hallucination-aware World Model-based Pol- icy Optimization (HaWMPO), a closed-loop reinforcement learning pipeline for VLA policy post-training with world models. Specifically, HaWMPO introduces an action-conditioned hallucination-aware model to estimate the reliability of gen- erated image sequences, and incorporates hallucination scores into group relative policy optimization through a Reward-Soft mechanism, suppressing unreliable ac- tion chunks during training. On the LIBERO benchmark, HaWMPO achieves the best average success rate, with gains of 15.0% over the base model and 2.8% over the strongest baseline; real-world experiments on a G1 robot further validate its effectiveness, raising the average success rate on two manipulation tasks from 67.5% to 80.0%.