arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2604.24764cs.CV

World-R1:通过强化学习为文本到视频生成注入3D约束

World-R1: Reinforcing 3D Constraints for Text-to-Video Generation

  • Zhejiang University(浙江大学)
  • Microsoft Research(微软研究院)
  • Independent Researcher(独立研究者)

机构由 AI 辅助整理,请以论文原文为准。

Weijie Wang, Xiaoxuan He, Youping Gu, Yifan Yang, Zeyu Zhang, Yefei He, Yanbo Ding, Xirui Hu, Donny Y. Chen, Zhiyuan He, Yuqing Yang, Bohan Zhuang

更新

AI总结:

提出World-R1框架,利用强化学习(Flow-GRPO)结合3D基础模型和视觉语言模型的反馈,在不修改架构的情况下增强视频生成的3D一致性,并采用周期解耦训练策略平衡刚体几何与动态场景。

AI中文摘要:

最近的视频基础模型展示了令人印象深刻的视觉合成能力,但经常遭受几何不一致性的困扰。现有方法尝试通过架构修改注入3D先验,但往往导致高计算成本并限制可扩展性。我们提出World-R1,一个通过强化学习将视频生成与3D约束对齐的框架。为促进这种对齐,我们引入了一个专门为世界模拟定制的纯文本数据集。利用Flow-GRPO,我们使用预训练的3D基础模型和视觉语言模型的反馈来优化模型,在不改变底层架构的情况下强制执行结构一致性。我们进一步采用周期解耦训练策略来平衡刚体几何一致性与动态场景流畅性。大量评估表明,我们的方法显著增强了3D一致性,同时保留了基础模型的原始视觉质量,有效弥合了视频生成与可扩展世界模拟之间的差距。

英文摘要:

Recent video foundation models demonstrate impressive visual synthesis but frequently suffer from geometric inconsistencies. While existing methods attempt to inject 3D priors via architectural modifications, they often incur high computational costs and limit scalability. We propose World-R1, a framework that aligns video generation with 3D constraints through reinforcement learning. To facilitate this alignment, we introduce a specialized pure text dataset tailored for world simulation. Utilizing Flow-GRPO, we optimize the model using feedback from pre-trained 3D foundation models and vision-language models to enforce structural coherence without altering the underlying architecture. We further employ a periodic decoupled training strategy to balance rigid geometric consistency with dynamic scene fluidity. Extensive evaluations reveal that our approach significantly enhances 3D consistency while preserving the original visual quality of the foundation model, effectively bridging the gap between video generation and scalable world simulation.

补充信息

↑