arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

PAVXploreRL:具有动作探索的物理动作视觉世界模型强化学习

PAVXploreRL: Physical-Action-Visual World Model Reinforcement Learning with Action Exploration

Han Wang, Zijun Wang, Shuoshuo Xue, Rui Cao, Fengjiao Chen, Xiaodan Liang, Roy Ka-Wei Lee

arXiv 2607.16602首次发表:更新:

发表机构

Singapore University of Technology and Design; Sun Yat-sen University(新加坡科技设计大学; 中山大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究针对动作条件世界模型泛化性差的问题,提出PAVXploreRL强化学习框架,基于预训练潜在世界模型,通过奖励驱动训练优化PAV目标,联合利用ID轨迹与OOD动作探索提升泛化能力,实验证明该方法性能更优。

AI 中文摘要

动作条件世界模型是具身人工智能的关键组成部分,可作为可扩展的策略评估器,减少对昂贵的现实世界部署的依赖。为了准确捕捉各种动作引起的动态,此类模型应满足物理合理性(P)、动作一致性(A)和视觉保真度(V)这三个关键目标,统称为PAV,同时对分布内(ID)专家演示和分布外(OOD)动作保持鲁棒性。然而,现有方法主要依赖ID动作视频对和像素级重建损失,未明确优化PAV目标,在专家数据之外泛化性较差。为解决此问题,我们提出PAVXploreRL,这是一个基于预训练潜在世界模型的强化学习框架,通过奖励驱动训练明确优化PAV目标。为提高动作泛化能力,我们的方法联合利用ID轨迹和噪声驱动的OOD动作探索,无需配对视频监督。实验表明,PAVXploreRL始终优于预训练基线,在基准测试中平均增益5.6%,并产生更高质量的PAV属性。作为策略评估器,它还能产生更可靠的性能估计,减少如Ctrl-World等仅基于专家的先前世界模型的高估偏差。

英文摘要

Action-conditioned world models are a key component of embodied AI, serving as scalable policy evaluators that reduce reliance on expensive real-world rollouts. To accurately capture diverse action-induced dynamics, such models should satisfy three key objectives-Physical Plausibility (P), Action Adherence (A), and Visual Fidelity (V), collectively referred to as PAV-while remaining robust to both in-distribution (ID) expert demonstrations and out-of-distribution (OOD) actions. However, existing methods primarily rely on ID action-video pairs and pixel-level reconstruction losses, which do not explicitly optimize PAV objectives and generalize poorly beyond expert data. To address this, we propose PAVXploreRL, a reinforcement learning framework built on a pretrained latent world model that explicitly optimizes PAV objectives through reward-driven training. To improve action generalization, our method jointly leverages ID trajectories and noise-driven OOD action exploration, without paired video supervision. Experiments show that PAVXploreRL consistently outperforms pretrained baselines, achieving a 5.6% average gain across benchmarks and producing higher-quality PAV properties. As a policy evaluator, it also yields more reliable performance estimates and reduces the overestimation bias of prior expert-only world models such as Ctrl-World. Code: https://github.com/Social-AI-Studio/PAVXploreRL

Comments9 pages, 5 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑