发表机构
Nanjing University; Jiutian Research(南京大学; 中移九天)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出VideoX-Qwen框架,通过大规模数据生产流水线生成120万条编辑记录,并采用统一Qwen-Wan编辑器,在11项指标中9项领先,实现高性能指令驱动视频编辑。
AI 中文摘要
通用视频编辑的进展依赖于构建大规模成对监督数据,并有效将视频生成骨干网络适配到指令驱动的编辑任务中。与视频生成不同,视频编辑必须执行所请求的变换,同时保留无关主体、场景结构、运动和时间连续性。我们提出了VideoX-Qwen,一个用于通用基于指令的视频编辑的集成数据构建与模型训练框架。我们的可扩展生产流水线将专门的生成和理解模型组织为互补路线,用于添加、移除、替换和属性编辑,随后进行质量筛选和指令丰富。该流水线生成了超过120万条定向视频编辑记录,每个主要任务组包含超过40万条记录,自动接受率达89%。所生成的语料库通过统一的源-指令-目标接口,提供了对常见编辑操作的广泛且结构化的覆盖。我们进一步开发了一个统一的Qwen-Wan编辑器,将多模态语义条件与密集源视频潜在引导相结合。一种渐进式图像-视频训练策略对齐了多模态指令接口,使视频生成器适应源条件编辑,并利用选定的高分辨率数据优化输出质量。在与UniVideo和Kling O1进行的100个示例对比中,VideoX-Qwen在11项报告指标中的9项上取得了最佳平均结果,包括指令遵循、编辑质量、内容保留、结构和感知相似性以及视频分布质量。总之,大规模数据生产系统和统一训练框架为更强大的指令驱动视频编辑提供了实用基础。
英文摘要
Progress in general-purpose video editing depends on constructing large-scale paired supervision and effectively adapting video-generation backbones to instruction-driven editing. Unlike video generation, video editing must execute a requested transformation while preserving unrelated subjects, scene structure, motion, and temporal continuity. We present VideoX-Qwen, an integrated data-construction and model-training framework for general instruction-based video editing. Our scalable production pipeline organizes specialized generation and understanding models into complementary routes for addition, removal, replacement, and attribute editing, followed by quality screening and instruction enrichment. It produces more than 1.2 million directional video-editing records, including over 400,000 records in each major task group, with an automatic acceptance rate of 89%. The resulting corpus provides broad and structured coverage of common editing operations through a unified source-instruction-target interface. We further develop a unified Qwen-Wan editor that combines multimodal semantic conditioning with dense source-video latent guidance. A progressive image-video training strategy aligns the multimodal instruction interface, adapts the video generator to source-conditioned editing, and refines output quality with selected high-resolution data. In a 100-example comparison with UniVideo and Kling O1, VideoX-Qwen achieves the best mean result on nine of eleven reported metrics, including instruction following, editing quality, content preservation, structural and perceptual similarity, and video-distribution quality. Together, the large-scale data-production system and unified training framework provide a practical foundation for more capable instruction-driven video editing.
CommentsTechnical report