结合世界模型的Q学习
Q-Learning With World Models
- Stanford University(斯坦福大学)
- Peking University(北京大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本研究提出QWM框架,将世界模型与标准Q学习结合,在真实环境训练中避免复合模型偏差,在Robomimic和LIBERO基准上的样本效率与性能均优于现有SOTA方法。
AI中文摘要:
离线强化学习(RL)已在样本效率上取得显著进展,支持将视觉-语言-动作模型进行RL微调,得到可靠且高性能的策略。世界模型可通过预测状态变化而非仅动作,进一步提升样本效率,但其成功大多局限于监督式策略学习。先前的基于模型的RL方法常直接在想象的回滚轨迹上优化策略或价值函数,易产生复合偏差,且难以扩展到真实机器人等大型高维问题,该问题会随任务时长和视觉复杂度增加而恶化。本研究探究能否在标准Q学习基础上直接利用世界模型,在保持在线真实环境训练的同时提升性能。我们提出QWM框架,该框架利用世界模型在Q学习基础上对想象轨迹执行测试时搜索,以便在线回滚和评估期间选择高价值动作。由于策略和价值函数仅在真实转换上训练,QWM避免了复合模型偏差,同时仍能从预测搜索中获得样本效率优势。在具有挑战性的操作基准Robomimic和LIBERO上,QWM在样本效率和性能方面均显著优于先前的强SOTA方法。
英文摘要:
Off-policy reinforcement learning (RL) has become increasingly sample-efficient, enabling applications such as RL fine-tuning of Vision-Language-Action models into reliable, high-performing policies. World models offer a further lever for sample efficiency, as they predict state changes rather than actions alone, but their success has largely been confined to supervised policy learning. Prior model-based RL methods often optimize the policy or value function directly on imagined rollouts, which is prone to compounding bias and struggles to scale to large, high-dimensional problems such as real-world robotics, a problem that worsens with task horizon and visual complexity. In this work, we instead ask whether we can leverage world models directly on top of standard Q-learning to improve performance, while remaining trained and grounded in the real, online setting. We propose QWM, a framework that leverages world models to perform test-time search over imagined trajectories on top of Q-learning to select high-value actions during both online rollouts and evaluation. Since the policy and value function are trained only on real transitions, QWM avoids compounding model bias while still gaining the sample-efficiency benefits of predictive search. On challenging manipulation benchmarks Robomimic and LIBERO, QWM significantly outperforms strong prior state-of-the-art methods on both sample efficiency and performance.