arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

将图像编辑器转换为视频编辑器

Transforming Image Editors into Video Editors

Feng Wang, Zijie Li, Ceyuan Yang, Alan Yuille, Peng Wang

arXiv 2610.11037首次发表:更新:

发表机构

ByteDance Seed; Johns Hopkins University(字节跳动Seed; 约翰斯·霍普金斯大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究提出基于锚点的视频编辑框架AVE,通过将强大图像编辑器转换为视频编辑器,在IVEBench和VIE-Bench上实现了指令遵循、时间一致性等方面的良好性能,为视频编辑提供了低成本替代方案。

AI 中文摘要

近期的图像编辑系统已在语义理解、视觉保真度和指令遵循能力方面取得了令人瞩目的成果,而视频编辑仍然难度大、成本高得多。在本文中,我们提出了一种端到端视频编辑的简单替代方案:我们不是训练一个整体式视频编辑器,而是通过基于锚点的生成将强大的图像编辑器转换为视频编辑器。我们的核心见解是,视频编辑可分解为两个子问题:编辑稀疏的关键帧集,并将这些编辑跨时间传播。基于这一观察,我们提出了基于锚点的视频编辑(Anchor-based Video Editing, AVE),这是一个两阶段框架:强大的图像编辑器首先对选定的关键帧执行组合编辑,然后运动引导的图像到视频扩散模型将编辑后的关键帧视为固定锚点,生成最终视频。这种设计直接继承了现代图像编辑器的优势,同时避免了昂贵的端到端视频编辑训练。在 IVEBench 和 VIE-Bench 上的实验表明,AVE 在指令遵循、时间一致性和内容保真度方面取得了强劲性能。进一步的 ablation 研究显示,最终视频编辑质量与图像编辑器的质量密切相关,这表明视频编辑的未来进展可能来自更强大的图像编辑基础以及向视频的轻量迁移。代码可在 https URL 获取。

英文摘要

Recent image editing systems have achieved impressive semantic understanding, visual fidelity, and instruction-following ability, while video editing remains substantially more difficult and costly. In this paper, we present a simple alternative to end-to-end video editing: instead of training a monolithic video editor, we transform a strong image editor into a video editor through anchor-based generation. Our key insight is that video editing can be decomposed into two subproblems: editing a sparse set of keyframes and propagating those edits across time. Based on this observation, we propose Anchor-based Video Editing (AVE), a two-stage framework in which a powerful image editor first performs composed editing on selected keyframes, and a motion-guided image-to-video diffusion model then generates the final video by treating the edited keyframes as fixed anchors. This design directly inherits the strengths of modern image editors while avoiding expensive end-to-end video editing training. Experiments on IVEBench and VIE-Bench show that AVE achieves strong performance in instruction following, temporal consistency, and content fidelity. Further ablations reveal that final video editing quality is strongly correlated with the quality of the image editor, suggesting that future progress in video editing may come from stronger image editing foundations and lightweight transfer to video. Code is available at https://github.com/wangf3014/AVE.

CommentsIn NeurIPS 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑