发表机构
The Chinese University of Hong Kong; Zhejiang University; ByteDance; The Ohio State University(香港中文大学; 浙江大学; 字节跳动; 俄亥俄州立大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
ThinkV2V提出推理驱动的视频编辑框架,通过MLLM-to-DiT架构显式激活思考,结合渐进式课程训练与推理时扩展,在5B模型上超越10B基线,并发布数据集与基准。
AI 中文摘要
指令引导的视频编辑已取得显著进展,然而现有方法主要将多模态大语言模型(MLLMs)用作语义编码器,因此在处理需要因果或语义推理的隐式编辑时常常力不从心。为弥合视频编辑中的这一根本性差距,我们提出ThinkV2V,一个面向复杂指令引导视频编辑的推理驱动框架,在视觉生成之前显式激活MLLM的思考过程。其核心在于,ThinkV2V构建于一种实用的MLLM到DiT架构之上,将对源视频和指令的显式思考转化为精细化的条件信号,用于视频编辑。此外,我们为其配备了专门的训练与推理方案,结合渐进式课程训练(Progressive Curriculum Training,逐步将模型从基础编辑培养至推理密集型案例)与推理时思考扩展(Inference-Time Thinking Scaling,迭代细化候选提示并选择最可靠的一个),以在具有挑战性的编辑场景中更好地激发推理能力。我们还整理了ThinkV2V-150K数据集,并引入ThinkV2V-Bench,以支持具有隐式意图和因果推理的视频编辑的训练与评估。实验结果表明,ThinkV2V在复杂和标准编辑场景下均达到最先进性能,其中我们5B规模的DiT模型显著优于更大的10B规模基线模型。
英文摘要
Instruction-guided video editing has made significant progress, yet existing methods use multimodal large language models (MLLMs) primarily as semantic encoders, so they often fall short in working with implicit edits that require causal or semantic reasoning. To bridge this fundamental gap in video editing, we propose ThinkV2V, a reasoning-driven framework for complex instruction-guided video editing, explicitly activating MLLM thinking before visual generation. At its core, ThinkV2V builds on a practical MLLM-to-DiT architecture to turn explicit thinking over the source video and instruction into refined conditioning signals for video editing. Further, we equip it with a dedicated training and inference recipe, combining Progressive Curriculum Training, which gradually cultivates the model from basic editing to reasoning-intensive cases, with Inference-Time Thinking Scaling, which iteratively refines candidate prompts and selects the most reliable one, to better elicit reasoning in challenging editing scenarios. We also curate the ThinkV2V-150K dataset and introduce ThinkV2V-Bench to support training and evaluation of video editing with implicit intent and causal reasoning. Experimental results demonstrate the state-of-the-art performance of ThinkV2V on both complex and standard editing scenarios, in which our 5B-scale DiT model substantially outperforms larger 10B-scale baselines.
CommentsProject page: https://correr-zhou.github.io/ThinkV2V