arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

Qwen-Video-Edit:通过复用图像编辑模型实现基于指令的视频编辑

Instruction-Based Video Editing by Repurposing an Image Editing Model

Yunpeng Bai, Yossi Gandelsman, Michaël Gharbi, Qixing Huang

arXiv 2608.14790首次发表:更新:

发表机构

UT Austin; Reve(德克萨斯大学奥斯汀分校; Reve公司)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究提出Qwen-Video-Edit,通过复用Qwen-Image-Edit图像编辑模型,经少量适配实现基于指令的视频编辑,证明图像编辑先验可迁移至视频编辑任务。

AI 中文摘要

基于指令的视频编辑通常构建在视频预训练的生成主干之上:需耗费大量成本对视频扩散变换器进行适配,使其能以源视频和编辑指令作为条件。本报告探索了一条不同的路径,证明强大的基于指令的图像编辑模型可通过直接在视频VAE隐空间上操作来编辑视频。我们从Qwen-Image-Edit出发,将Wan 2.1视频VAE的隐帧排列为一张大型虚拟图像的图块,复用编辑器的图像位置编码处理每个图块,并通过一对轻量型输入/输出投影层连接两个隐空间,该投影层从编辑器自身的patchify和unpatchify层进行热启动,使得在初始化时,(静态)视频能被模型精确嵌入为其已理解的图像。整个系统随后在公开的Ditto-1M编辑三元组上进行微调,Wan 2.2的少量去噪步骤可作为可选的时间增强器。我们通过一系列零训练观察结果为该设计提供支撑:原始图像编辑器可编辑呈现为联系表的视频;它对该表的token是来自联合编码还是隐空间中拼接的逐帧编码并无偏好;它甚至能零样本编辑真实视频隐空间至清晰可识别的程度,仅需通过微调弥合保真度差距即可。我们的结果表明,尽管在视频隐空间的训练上投入了大量资源,但逐帧视频隐空间仍与图像域足够接近,成熟的图像编辑先验仅需极少适配即可迁移。项目页面:this https URL;代码:this https URL;模型:this https URL。

英文摘要

Instruction-based video editing is commonly built on video-pretrained generative backbones: a video diffusion transformer is adapted, at considerable cost, to condition on a source video and an editing instruction. In this report we explore a different route and show that a strong instruction-based image editing model can edit videos by operating directly on video-VAE latents. Starting from Qwen-Image-Edit, we arrange the latent frames of a Wan~2.1 video VAE as tiles of one large virtual image, reuse the editor's image positional encoding for every tile, and bridge the two latent spaces with a pair of lightweight input/output projections warm-started from the editor's own patchify and unpatchify layers, so that at initialization a (static) video is embedded exactly as an image the model already understands. The whole system is then fine-tuned on the public Ditto-1M editing triplets, and a few denoising steps of Wan~2.2 serve as an optional temporal enhancer. We motivate the design with a chain of zero-training observations: the stock image editor already edits a video presented as a contact sheet; it is indifferent to whether the sheet's tokens come from one joint encode or from per-frame encodes stitched in latent space; and it even edits genuine video latents zero-shot to a clearly recognizable degree, leaving fine-tuning only a fidelity gap to close. Our results suggest that, despite the large investment in training video latent spaces, per-frame video latents remain close enough to the image domain that mature image editing priors transfer with minimal adaptation. Project Page: https://yunpeng1998.github.io/Qwen-Video-Edit-Page Code: https://github.com/yunpeng1998/Qwen-Video-Edit Model: https://huggingface.co/yunpeng1998/Qwen-Video-Edit

CommentsProject Page: https://yunpeng1998.github.io/Qwen-Video-Edit-Page Code: https://github.com/yunpeng1998/Qwen-Video-Edit Model: https://huggingface.co/yunpeng1998/Qwen-Video-Edit

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑