arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

WCM:用于视觉-语言-动作强化学习的世界评论者模型

WCM: A World Critic Model for Vision-Language-Action Reinforcement Learning

Senyu Fei, Xiaopeng Yu, Siyin Wang, Xianzhong Zhao, Jingjing Gong, Xipeng Qiu

arXiv 2607.29613首次发表:更新:

发表机构

Tongji University; Shanghai Innovation Institute; Fudan University(同济大学; 上海创新研究院; 复旦大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究针对视觉-语言-动作强化学习中评论者与机器人部分可观测性不匹配的问题,提出WCM模型,联合预测未来潜态与值估计,在149个仿真任务和7个真实任务上均实现最优性能与泛化性。

AI 中文摘要

视觉-语言-动作(VLA)模型的强化学习(RL)后训练在机器人操纵领域展现出强大潜力。在RL方法中,基于评论者的方法依赖于主要基于单帧观测或单帧VLM骨干潜变量运行的值估计器,这与机器人控制的部分可观测性存在根本不匹配。将观测历史纳入评论者的朴素方法会在高维视觉空间中产生指数级复杂度,且仍会失败,因为纯标量回报回归无法为跨时间动态的学习提供足够监督。我们将根本原因识别为状态近似问题:若无显式世界建模目标,评论者的表示无法捕捉准确值估计所需的时间结构。为解决此问题,我们提出基于轻量LeJEPA架构构建的世界评论者模型(WCM);WCM联合预测未来潜状态并估计值,使评论者的表示被显式训练以捕捉时间动态,而非仅回归标量回报。WCM可无缝集成到同策略和异策略训练流程,且与包括Pi0、Pi0.5和OpenVLA-OFT在内的最先进VLA骨干兼容。在四个基准的149个任务上开展的大量实验表明,WCM在分布内和分布外设置中均持续实现最先进性能,且泛化增益尤为显著。我们进一步使用OpenVLA-OFT和Pi0.5在七个真实世界操纵任务上通过异策略RL验证WCM,确认其可在多样设置中稳定部署。

英文摘要

Reinforcement learning (RL) post-training of Vision-Language-Action (VLA) models has shown strong promise for robotic manipulation. Among RL methods, critic-based approaches rely on a value estimator that predominantly operates on single-frame observations or single-frame VLM backbone latents, which is a fundamental mismatch with the partially observable nature of robot control. A naive approach to incorporate observation history into the critic incurs exponential complexity with high-dimensional visual space, and still fails because pure scalar-return regression provides insufficient supervision for learning cross-temporal dynamics. We identify the root cause as a state approximation problem: without an explicit world modeling objective, the critic's representation cannot capture the temporal structure needed for accurate value estimation. To address this, we propose the World Critic Model (WCM), built on a lightweight LeJEPA architecture; WCM jointly predicts future latent state and estimates values, such that the critic's representation is explicitly trained to capture temporal dynamics rather than merely regress scalar returns. WCM integrates seamlessly into both on-policy and off-policy training pipelines and is compatible with state-of-the-art VLA backbones including Pi0, Pi0.5, and OpenVLA-OFT. Extensive experiments on 149 tasks across four benchmarks demonstrate that WCM consistently achieves state-of-the-art performance in both in-distribution and out-of-distribution settings, with particularly strong generalization gains. We further validate WCM on seven real-world manipulation tasks using OpenVLA-OFT and Pi0.5 with off-policy RL, confirming stable deployment across diverse settings.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑