RefVideo-6M:一个用于教学视频编辑的可靠基于参考的数据集
RefVideo-6M: A Reliable Reference-Based Dataset for Instructional Video Editing
浏览论文内容
中文总结 AI 辅助
本文针对现有视频编辑数据集的局限,构建含600万样本的RefVideo-6M参考引导编辑数据集,训练Ref-MoT模型,实验证明其监督更可靠,可提升编辑模型的视觉质量、可控性与参考一致性。
中文摘要 AI 辅助
视频编辑领域的近期进展在很大程度上由大规模基于指令的数据集推动。然而,现有数据集仍存在两个关键局限:其一,目标视频通常由自动编辑模型生成,可能引入明显的人工痕迹和不可靠的监督信号;其二,多数公开数据集主要依赖文本指令,缺乏对精确、身份保留及可控编辑至关重要的视觉参考。为解决这些局限,本文提出RefVideo-6M,这是一个大规模参考引导编辑数据集,包含500万个视频编辑样本和100万个图像编辑样本。为确保监督的可靠性,该数据集采用构建流程,将无人工痕迹的真实视频作为编辑目标,并借助多位编辑专家生成经质量筛选的输入条件。此外,它提供约600万个视觉参考,涵盖多样的参考类型和编辑场景,使模型能够学习超越纯文本指令的细粒度视觉对应关系。基于RefVideo-6M,本文进一步训练了一个参考引导视频编辑模型Ref-MoT,以评估所提出数据集的有效性和可扩展性。大量实验表明,RefVideo-6M比现有数据集提供显著更可靠的监督,且能训练出视觉质量、可控性和参考一致性均有所提升的强大编辑模型。该开源数据集可在指定网址获取。
英文摘要
Recent advances in video editing have been largely driven by large-scale instruction-based datasets. However, existing datasets still suffer from two critical limitations. First, target videos are commonly produced by automatic editing models, which may introduce visible artifacts and unreliable supervision signals. Second, most public datasets rely primarily on textual instructions, while lacking visual references that are crucial for precise, identity-preserving, and controllable editing. To address these limitations, we introduce RefVideo-6M, a large-scale reference-guided editing dataset containing 5 million video editing samples and 1 million image editing samples. To ensure reliable supervision, our dataset uses a construction pipeline that treats artifact-free real videos as editing targets and generates quality-filtered input conditions with multiple editing experts. In addition, it provides approximately 6 million visual references, covering diverse reference types and editing scenarios, thereby enabling models to learn fine-grained visual correspondence beyond text-only instructions. Based on RefVideo-6M, we further train a reference-guided video editing model, Ref-MoT, to evaluate the effectiveness and scalability of the proposed dataset. Extensive experiments demonstrate that RefVideo-6M provides substantially more reliable supervision than existing datasets and enables the training of powerful editing models with improved visual quality, controllability, and reference consistency. The open-source dataset is available at https://huggingface.co/datasets/RefVideo6M/RefVideo6M.
发表机构
- Institute of Artificial Intelligence, China Telecom (TeleAI)(中国电信人工智能研究院(TeleAI))
- The Chinese University of Hong Kong (CUHK)(香港中文大学)
- Sun Yat-sen University(中山大学)
- Fudan University(复旦大学)
- Tsinghua University(清华大学)
机构由 AI 辅助整理,请以论文原文为准。