PWM:基于在线强化学习的个性化世界模型
PWM: Personalized World Models with Online Reinforcement Learning
浏览论文内容
中文总结 AI 辅助
本文提出PWM框架,利用在线强化学习从短视频定制世界模型,通过GRPO或DiffusionNFT优化LoRA适配器,在PWM-Bench的150个任务中显著提升场景定制质量。
中文摘要 AI 辅助
预训练的世界模型能够生成多样化的环境,但用户通常希望探索由其自身视频指定的特定场景。这要求在学习场景视觉特征的同时,保持动作条件生成的质量。我们提出了个性化世界模型(PWM),一种通过在线强化学习从短视频中定制交互式世界模型的框架。在PWM中,支持轨迹及其相关控制为从当前策略采样的延续提供奖励反馈。在GRPO实例化中,组相对优化使用统一的奖励(涵盖场景外观、视觉连续性和运动)更新紧凑的LoRA适配器,而基础策略锚定则正则化对冻结的Yume-5B骨干网络的预训练生成先验的更改。相同的适配过程应用于真实和渲染环境。我们还使用DiffusionNFT实例化PWM,作为学习场景特定适配器的替代奖励引导优化方法。我们还引入了PWM-Bench,包含室内、室外和游戏三个领域的150个定制任务,并在保留的延续上进行配对评估。PWM的GRPO和DiffusionNFT实例分别在71.3%和65.3%的评估场景中优于原生Yume,且在三个领域均获得正的平均增益。对于GRPO实例化,匹配的SFT比较进一步表明,在每个领域中获得更高的平均定制增益和更好的平均图像质量分数,同时保持接近预训练模型的帧级视觉质量。
英文摘要
Pretrained world models can generate diverse environments, yet users often want to explore a particular scene specified by their own video. This requires learning the scene's visual identity while retaining the quality of action-conditioned generation. We introduce Personalized World Models (PWM), a framework for customizing interactive world models from short scene videos through online reinforcement learning. In PWM, the support trajectory and its associated controls provide reward feedback on continuations sampled from the current policy. In the GRPO instantiation, group-relative optimization updates a compact LoRA adapter using a unified reward for scene appearance, visual continuity, and motion, while base-policy anchoring regularizes changes to the pretrained generation prior of a frozen Yume-5B backbone. The same adaptation procedure is applied across real and rendered environments. We also instantiate PWM with DiffusionNFT as an alternative reward-guided optimization method for learning the scene-specific adapter. We also introduce PWM-Bench, comprising 150 customization tasks across Indoor, Outdoor, and Gaming, with paired evaluation on held-out continuations. The GRPO and DiffusionNFT instantiations of PWM improve customization over native Yume in 71.3% and 65.3% of the evaluated scenes, respectively, with positive mean gains across all three domains. For the GRPO instantiation, matched SFT comparisons further demonstrate higher mean customization gains and better mean image-quality scores in every domain, while retaining frame-level visual quality close to the pretrained model.
发表机构
- Peking University(北京大学)
- La Trobe University(拉筹伯大学)
- Northwestern University(西北大学)
机构由 AI 辅助整理,请以论文原文为准。