arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

VINCIE-NExT:通过上下文建模实现从图像到视频的编辑

VINCIE-NExT: Unlocking Video Editing from Images via In-Context Modeling

Leigang Qu, Feng Cheng, Ziyan Yang, Bangbang Yang, Zhaoyang Huang, Wei Chow, Yicong Li, Wenjie Wang, Tat-Seng Chua, Yan Zeng

arXiv 2610.12104首次发表:更新:

发表机构

National University of Singapore; ByteDance Seed; University of Science and Technology of China(新加坡国立大学; 字节跳动种子项目; 中国科学技术大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究提出VINCIE-NExT框架,通过上下文视觉演示将图像编辑能力迁移至视频,分解视频编辑为子任务链,引入新型位置编码实现外观对齐,在OpenVE-Bench上达到当前最优编辑性能。

AI 中文摘要

构建一个实用的视频编辑器,其难度远高于视频生成器:编辑任务需要(源图像、指令、编辑后图像)三元组,这类数据的标注成本极高,且难以大规模合成;而图像编辑领域已发展成熟,拥有数百万对可用的编辑数据。本研究提出VINCIE-NExT,这是一个通过上下文视觉演示将编辑能力从图像迁移至视频的统一框架,无需依赖大规模配对的视频编辑数据。VINCIE-NExT将视频编辑分解为结构化的可组合子任务链(视频→图像→图像→视频),通过图像域传递编辑意图,支持在统一扩散目标下,从异构的图像和视频语料库进行可扩展的联合训练。模型会将用户提供或自身合成的图像编辑对作为上下文视觉演示,作为每个输出帧的空间外观蓝图。为在交错的上下文间实现外观编辑的对齐,本研究引入了一种新型位置编码,将图像演示与视频帧链接到共享的空间坐标系中,确保外观变化能以像素级精度传播到每个输出帧。编辑链还提供了原则性的测试时扩展方案:通过将子任务链作为渐进式扩散阶段执行,无需重新训练,仅投入额外计算资源即可提升编辑质量。在OpenVE-Bench上开展的全面实验表明,该模型在各类编辑类别中达到了当前最优性能, ablation实验也验证了各组件的有效性。

英文摘要

Building a capable video editor remains significantly harder than a video generator: editing requires (source, instruction, edited) triplets that are prohibitively expensive to annotate and difficult to synthesize at scale, whereas image editing has already reached maturity with millions of such pairs readily available. In this work, we introduce VINCIE-NExT, a unified framework that transfers editing capability from images to videos through in-context visual demonstrations, alleviating the need for large-scale paired video editing data. VINCIE-NExT decomposes video editing into a structured chain of composable sub-tasks (Video -> Image -> Image -> Video), routing editing intent through the image domain and enabling scalable joint training from heterogeneous image and video corpora under a unified diffusion objective. An image editing pair, synthesized by the model or supplied by the user, is prepended as an in-context visual demonstration that serves as a spatial appearance blueprint for every output frame. To ground appearance edits across the interleaved context, we introduce a novel position encoding that links image demonstrations and video frames in a shared spatial coordinate system, enabling pixel-faithful propagation of appearance changes to every output frame. Chain-of-Editing further provides principled test-time scaling: by executing the sub-task chain as progressive diffusion stages, editing quality can be improved by investing additional compute without retraining. Comprehensive experiments on OpenVE-Bench demonstrate the state-of-the-art performance across diverse editing categories, with ablations confirming the effectiveness of each component.

CommentsAccepted to NeurIPS'26. Project page: https://vincie-next.github.io/

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑