发表机构
Shanghai Jiao Tong University; Tencent Youtu Lab; Zhejiang University(上海交通大学; 腾讯优图实验室; 浙江大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对文本指令视频编辑难以传达细粒度视觉信息的问题,提出视觉上下文编辑新范式,构建VicEdit-400K数据集,开发VicEdit框架,在VicEditBench上实现SOTA性能。
AI 中文摘要
尽管基于指令的视频编辑已取得进展,但单模态文本指令在传达细粒度纹理和复杂动态方面存在固有困难。为弥合这一感知差距,我们提出视觉上下文编辑,这是一种将视频编辑从文本指令提升至涵盖单张图像、图像对和视频对的多模态视觉指导的新范式。为推动该范式发展,我们构建了首个用于视觉上下文视频编辑的大规模数据集VicEdit-400K,开发了自动化流水线以生成涵盖10种任务类型的40万个高质量样本,通过多维度过滤确保了卓越的视觉保真度和语义一致性。基于此基础,我们引入VicEdit,这是一个连接视觉与文本上下文的统一框架。为从异构参考中自适应提取编辑语义,我们设计了模态自适应语义蒸馏,可从视觉参考生成模态特定的语义标记,这些标记随后通过双上下文注入与文本指令协同整合,使生成过程能同时受益于视觉和文本信号。在VicEditBench上的广泛评估表明,VicEdit在基础指令编辑和视觉上下文编辑任务中均达到了SOTA性能,确立了视觉上下文学习是一种强大且可控的视频编辑范式。
英文摘要
Despite progress in instruction-based video editing, unimodal textual instructions inherently struggle to convey fine-grained textures and complex dynamics. To bridge this perceptual gap, we propose Visual In-context Editing, a new paradigm elevating video editing from textual instructions to multi-modal visual guidance encompassing single image, image pair, and video pair. To facilitate this paradigm, we curate VicEdit-400K, the first large-scale dataset for visual in-context video editing. We develop an automated pipeline to generate 400K high-quality samples across ten task types, ensuring superior visual fidelity and semantic consistency through multi-dimensional filtering. Leveraging this foundation, we introduce VicEdit, a unified framework to bridge visual and textual contexts. To adaptively extract editing semantics from heterogeneous references, we design Modality-Adaptive Semantic Distillation, which produces modality-specific semantic tokens from visual references. These tokens are then synergistically integrated with textual instructions through Dual-Context Injection, enabling the generation process to benefit from both visual and textual signals. Extensive evaluations on VicEditBench demonstrate that VicEdit achieves state-of-the-art performance across both basic instruction editing and visual in-context editing tasks, establishing visual in-context learning as a powerful and controllable paradigm for video editing.