arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2610.04152cs.CV

Kepler4D:通过4D场景状态演化实现可控的未来视频生成

Kepler4D: Controllable Future Video Generation via 4D Scene State Evolution

Feiran Wang, Bin Duan, Junyi Wu, Gaowen Liu, Yan Yan

首次发表
浏览论文内容

中文总结 AI 辅助

Kepler4D通过显式4D场景状态演化,利用Chain-of-Motion和视觉语言模型生成可控的未来视频,实现物体运动控制与场景一致性。

中文摘要 AI 辅助

视频世界模型旨在保持场景结构并预测动态物体在视觉观察之外的演化。我们提出了Kepler4D,一个通过显式4D场景状态演化进行未来视频生成的框架。给定单目视频,Kepler4D构建背景几何、物体运动历史、粗略空间支持和语义上下文的共享3D表示。Chain-of-Motion总结观察到的运动,并使用视觉语言模型选择结构化的速度和航向决策,并决定是否从下方限制物体中心高度。确定性展开将这些决策转化为未来的物体轨迹,以供在合成前进行检查和编辑。我们将演化的代理渲染为预训练视频生成器的几何控制,将粗略物体运动与外观和关节的合成分离。在真实世界视频上的实验表明,Kepler4D能够实现可控的物体运动和合理的未来展开,同时保持场景一致性。

英文摘要

Video world models aim to preserve scene structure and predict how dynamic objects evolve beyond visual observations. We present Kepler4D, a framework for future video generation through explicit 4D scene state evolution. Given a monocular video, Kepler4D constructs a shared 3D representation of background geometry, object motion histories, coarse spatial supports, and semantic context. Chain-of-Motion summarizes observed motion and uses a vision-language model to select structured speed and heading decisions and decide whether to bound object-center height from below. A deterministic rollout converts these decisions into future object trajectories for inspection and editing before synthesis. We render the evolving proxies into geometric controls for a pretrained video generator, separating coarse object motion from the synthesis of appearance and articulation. Experiments on real-world videos demonstrate that Kepler4D enables controllable object motion and plausible future rollout while preserving scene consistency.

发表机构

  • University of Illinois at Chicago(伊利诺伊大学芝加哥分校)
  • University of Michigan Ann Arbor(密歇根大学安娜堡分校)
  • Cisco(思科)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑