arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2610.12382cs.CV

WorldAlign:用于世界一致性视频生成的解耦4D奖励

WorldAlign: Decoupled 4D Reward for World-Consistent Video Generation

Jing He, Kaixin Ding, Xingye Tian, Guibao Shen, Wenhang Ge, Xin Tao, Pengfei Wan, Ying-Cong Chen

首次发表
浏览论文内容

中文总结 AI 辅助

本文提出解耦4D奖励框架WorldAlign,通过分离静态区域与动态主体适配对应世界先验,在Wan2.1和Wan2.2上同时提升视频的静态与动态一致性,无需人类偏好标注。

中文摘要 AI 辅助

忠实的视觉世界模拟要求生成的视频保持4D世界一致性,涵盖静态一致性与动态一致性。静态一致性要求静态环境在不同视角下具备连贯的3D结构,动态一致性则要求主体运动合理且外观随时间保持一致。几何感知的后训练是提升世界一致性的有效途径,但现有方法常依赖静态场景假设;即便适配动态场景的方法,也难以提供可靠的静态一致性反馈,且动态一致性常被忽视或评估不足。为解决这些局限,本文提出WorldAlign,一种解耦4D奖励框架,它在语义上分离静态区域与动态主体,并通过适配各自假设的世界先验为两者提供反馈。对于静态区域,WorldAlign通过语义引导的掩码重投影将静态几何与几何世界先验对齐,实现更可靠的静态一致性评估;辅助的相机运动奖励会抑制近乎静态的解决方案。对于动态主体,WorldAlign采用强大的视觉语言模型(VLM)作为动态世界先验,并基于样本特定的检查清单构建VLM作为评判者的奖励,该清单评估动态性、物理合理性、形状及纹理一致性。这种解耦设计支持无需人类偏好标注的更有效在线后训练。在两个预训练的图像到视频生成器Wan2.1和Wan2.2上,WorldAlign相比现有方法同时提升了静态与动态一致性,且未抑制整体或主体运动。这些结果表明,解耦的世界先验对齐可实现更忠实的视觉世界模拟。项目页面:this https URL

英文摘要

Faithful visual world simulation requires generated videos to maintain 4D world consistency, encompassing both static and dynamic consistency. Static consistency requires coherent 3D structure in static environments across viewpoints, while dynamic consistency requires plausible subject motion and consistent appearance over time. Geometry-aware post-training offers a promising way to improve world consistency. However, existing methods often rely on a static-scene assumption. Even those that accommodate dynamic scenes struggle to provide reliable static-consistency feedback, while dynamic consistency is often overlooked or inadequately assessed. To address these limitations, we introduce WorldAlign, a decoupled 4D reward framework that semantically separates static regions and dynamic subjects and provides feedback by aligning each with a world prior suited to its assumptions. For static regions, WorldAlign aligns static geometry with a geometric world prior through semantically guided masked reprojection, enabling more reliable static-consistency evaluation; an auxiliary camera-motion reward discourages nearly static solutions. For dynamic subjects, WorldAlign uses a strong vision-language model (VLM) as a dynamic world prior and constructs a VLM-as-a-judge reward based on sample-specific checklists that assess dynamicity, physical plausibility, shape, and texture consistency. This decoupled design enables more effective online post-training without requiring human preference annotations. Across two pretrained image-to-video generators, Wan2.1 and Wan2.2, WorldAlign jointly improves static and dynamic consistency over existing methods without suppressing overall or subject motion. These results support decoupled world-prior alignment for more faithful visual world simulation. Project page: https://worldalign.github.io/.

发表机构

  • The Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州))
  • The Hong Kong University of Science and Technology(香港科技大学)
  • KlingAI
  • The University of Hong Kong(香港大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑