优先回放用于高效基于世界模型的视觉-语言-动作策略优化
Prioritized Rollouts for Efficient World Model-based Vision-Language-Action Policy Optimization
- National Key Laboratory for Novel Software Technology, Nanjing University(南京大学计算机软件新技术全国重点实验室)
- School of Artificial Intelligence, Nanjing University(南京大学人工智能学院)
- Cirquar Technologies
- Mila – Quebec AI Institute(Mila – 魁北克人工智能研究所)
- Université de Montréal(蒙特利尔大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本文提出U-GROW,一种基于策略不确定性的轻量级采样层,通过优先回放信息丰富的状态,提升基于世界模型的VLA策略优化效率,并在模拟和真实操作任务中验证其有效性。
AI中文摘要:
视觉-语言-动作(VLA)模型已成为具身智能的强大范式,但使用强化学习(RL)对其进行微调仍受限于真实世界机器人交互的成本。基于模型的强化学习(MBRL)通过使用学习到的世界模型生成回放数据以进行策略优化,从而降低了这一成本。然而,随着VLA策略和世界模型的规模扩大,其计算成本变得高昂。现有方法通常平等对待所有状态,忽略了它们在策略改进效用上的显著差异。在本文中,我们表明策略不确定性有助于识别具有更大策略改进潜力的状态。策略仅在少量状态子集上表现出高不确定性,通常出现在决策敏感阶段,此时微小的动作差异可能改变任务结果,这表明在这些状态上的策略改进可能特别有价值。基于这些发现,我们引入了U-GROW,一个轻量级、即插即用的采样层,将更多模型回放引导至这些信息丰富的状态。通过仅修改分支起始分布,U-GROW可以集成到现有的MBRL流程中,而无需改变策略优化目标。在模拟和真实世界操作任务中的实验证明了U-GROW的效率和有效性,支持使用策略不确定性来指导经验生成。
英文摘要:
Vision-Language-Action (VLA) models have emerged as a powerful paradigm for embodied intelligence, but fine-tuning them with reinforcement learning (RL) remains constrained by the cost of real-world robot interaction. Model-based reinforcement learning (MBRL) reduces this cost by using a learned world model to generate rollouts for policy optimization. However, it becomes computationally expensive as VLA policies and world models scale. Existing methods typically treat states equally, overlooking substantial differences in their utility for policy improvement. In this paper, we show that policy uncertainty helps identify states with greater potential for policy improvement. The policy exhibits high uncertainty at only a small subset of states, often during decision-sensitive stages where small action differences can alter task outcomes, suggesting that policy improvements at these states could be particularly valuable. Building on these findings, we introduce U-GROW, a lightweight, plug-and-play sampling layer that directs more model rollouts to these informative states. By modifying only the branched-start distribution, U-GROW can be integrated into existing MBRL pipelines without changing the policy optimization objective. Experiments in both simulated and real-world manipulation tasks demonstrate the efficiency and effectiveness of U-GROW, supporting the use of policy uncertainty to guide experience generation.