发表机构
Wuhan University(武汉大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
提出MDN-Control,一种无需训练的多主体视频编辑框架,通过掩码定位、深度遮挡控制和噪声潜在提示,在MSVBench上取得最优编辑质量,解决属性泄漏与遮挡模糊问题。
AI 中文摘要
多主体视频编辑在修改指定主体的同时需保留非目标内容,但面临跨主体属性泄漏和遮挡模糊问题。现有方法依赖掩码,难以区分重叠主体或确保生成一致性。为解决这些局限,我们提出MDN-Control,一种无需训练的训练自由框架,联合控制目标定位、遮挡几何和外观初始化。具体而言,掩码引导的定位提供一致的目标定位,而深度感知的遮挡控制解决重叠主体间的模糊边界。我们进一步引入噪声潜在提示,从噪声库中检索高斯初始化以获得提示相关的先验。在MSVBench上的实验表明,MDN-Control实现了最低的CM-Err和最高的Q-Edit,同时保持竞争力的文本对齐和时间一致性,证明了结合空间、几何和潜在先验进行多主体视频编辑的有效性。
英文摘要
Multi subject video editing modifies designated subjects while preserving non target content, but faces cross subject attribute leakage, and occlusion ambiguity. Existing approaches rely on masks and struggle to distinguish overlapping subjects or ensure consistent generation. To address these limitations, we propose MDN-Control, a training free framework jointly controlling target localization, occlusion geometry, and appearance initialization. Specifically, mask-guided localization provides consistent target localization, while depth-aware occlusion control resolves ambiguous boundaries between overlapping subjects. We further introduce noise latent prompting, which retrieves Gaussian initializations from a noise library for prompt relevant priors. Experiments on MSVBench show that MDN-Control achieves the lowest CM-Err and the highest Q-Edit, while maintaining competitive text alignment and temporal consistency, demonstrating the effectiveness of combining spatial, geometric, and latent priors for multi subject video editing.
Comments5 pages, 3 figures