arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.38146cs.CV

LIFT:大视角变化下基于在线策略自蒸馏的未来布局视频生成

LIFT: Layout-In-Future Video Generation under Large Viewpoint Change via On-Policy Self-Distillation

Shengxiang Ji, Boyang Wang, Haiyang Xu, Bingnan Li, Yucheng Mao, Zeyuan Chen, Xiaojun Shan, Xiang Zhang, Gang Hua, Jianwen Xie, Zezhou Cheng, Zhuowen Tu

首次发表
浏览论文内容

中文总结 AI 辅助

LIFT提出统一框架,通过未来布局控制与在线策略自蒸馏,解决大视角变化下视频生成的内容与布局可控性问题,提升视频质量与可控性。

中文摘要 AI 辅助

我们提出了LIFT,一个统一的图像到视频生成框架,它将相机控制与未来布局(Layout-In-FuTure)控制相结合,使用户能够指定未来视图中应出现的内容及其位置。这解决了可控视频生成中的一个实际需求:给定初始图像,用户不仅关心相机如何移动,还关心场景在关键未来时刻(尤其是最后一帧)的外观。现有的相机控制指定了视角轨迹,而文本提示仅提供粗略的语义指导;两者都不能精确确定未来视图的内容和空间布局。这种限制在大视角变化下尤为明显,因为相机会揭示第一帧中不可见的区域。因此,LIFT将最后一帧的布局作为期望未来场景的显式控制信号。由于从这种稀疏布局指导中学习比基于密集逐帧布局的条件生成更具挑战性,我们引入了在线策略自蒸馏(OPSD),将密集布局教师模型的控制能力迁移到最后一帧布局学生模型。我们进一步构建了LIFT-Vista数据集,该数据集具有大视角变化以及相机和时间一致的布局标注。实验表明,LIFT在视频质量、未来布局可控性和相机可控性方面优于其他方法。

英文摘要

We introduce LIFT, a unified image-to-video generation framework that complements camera control with Layout-In-FuTure control, enabling users to specify what should appear in a future view and where it should appear. This addresses a practical need in controllable video generation: given an initial image, users often care not only about how the camera moves, but also about what the scene should look like at key future moments, especially the final frame. Existing camera controls specify viewpoint trajectories, while text prompts provide only coarse semantic guidance; neither precisely determines the content and spatial layout of future views. This limitation becomes particularly pronounced under large viewpoint changes, where the camera reveals regions that are not visible in the first frame. LIFT therefore uses the last-frame layout as an explicit control signal for the desired future scene. Since learning from such sparse layout guidance is substantially more challenging than conditioning on dense per-frame layouts, we introduce on-policy self-distillation (OPSD) to transfer the control capability of a dense-layout teacher to a last-frame-layout student. We further curate LIFT-Vista, a dataset featuring large viewpoint changes with camera and temporally consistent layout annotations. Experiments show that LIFT improves video quality, future-layout controllability, and camera controllability over other methods.

发表机构

  • UC San Diego(加州大学圣迭戈分校)
  • University of Virginia(弗吉尼亚大学)
  • Meta
  • Amazon(亚马逊)
  • Lambda

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑