发表机构
University of Science and Technology of China; Zhongguancun Academy; Tsinghua University(中国科学技术大学; 中关村学院; 清华大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出CoT-Edit的plan--guide--edit框架,以CoT增强的MLLM为规划器生成空间先验,结合扩散编辑器实现高保真指令视频编辑,性能优于多个基准方法
AI 中文摘要
复杂场景下基于文本驱动的指令视频编辑仍具挑战性:纯文本提示常无法捕捉精确的空间关系与物理约束,导致目标模糊、生成结果不符合物理规律。为解决该问题,本文提出plan--guide--edit框架,明确搭建语义意图与空间执行的桥梁。该框架中,思维链(CoT)增强的多模态大语言模型(MLLM)作为规划器,对视频与指令进行结构化推理,生成精确的边界框序列及带属性的编辑指令。这些空间先验随后指导框条件掩码生成器,将模糊的全局检索转化为局部化、上下文感知的优化,生成能更准确捕捉物体尺度、接触关系与位置的掩码。基于这些空间与语义信号,基于扩散的编辑器整合掩码、带属性的指令及帧特征,渲染出时间连贯、空间对齐良好的高保真编辑结果。该框架先以模块化方式训练,再进行联合训练,在降低数据需求的同时实现了更优性能,在含多个相似物体的场景中实现了精确定位,且能生成物理上一致的物体添加效果;大量实验表明,该框架在多个强基准方法上达到了最先进性能。更多细节可访问:this https URL
英文摘要
Text-driven instruction-based video editing in complex scenes remains challenging: purely textual prompts often fail to capture precise spatial relationships and physical constraints, resulting in target ambiguity and physically implausible outcomes. To address this, we propose a plan--guide--edit framework that explicitly bridges semantic intent and spatial execution. In our framework, a Chain-of-Thought (CoT)-enhanced multimodal large language model (MLLM) serves as a planner, performing structured reasoning over the video and instructions to derive a precise sequence of bounding boxes and attribute-enriched editing directives. These spatial priors then guide a box-conditioned mask generator, transforming ambiguous global retrieval into localized, context-aware refinement and producing masks that more accurately capture object scale, contact relationships, and placement. Building on these spatial and semantic signals, a diffusion-based editor integrates the masks, enriched instructions, and frame features to render high-fidelity edits that remain temporally coherent and spatially well aligned. Trained first in a modular manner and then jointly, our framework achieves superior performance with reduced data requirements, delivering precise localization in scenes with multiple similar objects and physically consistent object additions, and extensive experiments demonstrate state-of-the-art performance over multiple strong baseline methods. More details are available at: https://github.com/flying-sky999/CoT-Edit