arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.36496cs.CV

重新构想视频动态

Reimagine Video Dynamics

Yu Yuan, Yawen Lu, Guoxian Song, Kevin Duarte, Ratheesh Kalarot, Di Chang, Xijun Wang, Stanley H. Chan

首次发表
浏览论文内容

中文总结 AI 辅助

提出RVD框架,通过自监督重建分离动态令牌,实现语言引导的视频动态编辑、免训练重定时和外观控制重渲染。

中文摘要 AI 辅助

大多数视频编辑方法侧重于改变源视频的外观,而对动态的控制能力有限。我们提出了重新构想视频动态(RVD)框架,该框架将紧凑、可编辑的动态令牌从视觉上下文中分离出来。我们通过自监督重建学习该令牌:给定第一帧作为视觉上下文,渲染器必须从动态令牌恢复原始视频,从而鼓励令牌捕捉场景如何演变而非外观。这种分离使得视频动态能够直接编辑,同时保留视觉上下文。我们开发了一个语言引导的动态令牌编辑器,将源动态转换为目标动态,并通过可扩展的反事实视频对流水线和两阶段训练策略进行训练。大量实验表明,RVD能够实现有效的视频动态编辑、免训练的重定时以及外观控制的重新渲染。

英文摘要

Most video editing methods focus on changing the appearance of the source video, while offering limited control over its dynamics. We introduce Reimagine Video Dynamics (RVD), a framework that disentangles a compact, editable dynamics token from visual context. We learn this token through self-supervised reconstruction: given the first frame as visual context, a renderer must recover the original video from the dynamics token, encouraging it to capture how the scene evolves rather than how it looks. This disentanglement allows video dynamics to be edited directly while preserving visual context. We develop a language-guided dynamics-token editor that transforms source dynamics into target dynamics, and train it with a scalable counterfactual video-pair pipeline and a two-stage training strategy. Extensive experiments show that RVD enables effective video dynamics editing, training-free retiming, and appearance-controlled re-rendering.

发表机构

  • Adobe
  • Purdue University(普渡大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑