从带草稿到无草稿:通过特权蒸馏和快速植入实现一步视频对象移除
From Draft to Draft-Free: One-Step Video Object Removal via Privileged Distillation and Fast Planting
- Institute of Artificial Intelligence and Robotics, Xi’an Jiaotong University(西安交通大学人工智能与机器人研究所)
- SGIT AI Lab, State Grid Corporation of China(国家电网公司SGIT人工智能实验室)
- Zhejiang University of Technology(浙江工业大学)
- Baidu(百度)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
研究视频对象移除问题,提出D2DF框架,通过特权蒸馏和自引导快速植入模块,将教师模型多步细化能力提炼到学生模型,实现一步视频对象移除,在质量和效率上优于传统及多步生成方法。
AI中文摘要:
视频对象移除是视频编辑中一项基本但具有挑战性的任务。尽管近期有进展,但现有方法分为两类。传统方法常引入明显伪影且结果不自然,基于扩散的方法视觉效果好但需多步去噪,实用性受限。我们提出D2DF框架,将把粗糙草稿转化为精细视频的能力提炼到一步视频生成模型中。训练教师模型将低质量移除结果细化为高保真视频,通过PPCD将此能力提炼到学生模型。引入SGFP模块消除对草稿的依赖,实现完全无草稿的一步模型。实验表明,有草稿和无草稿版本在多个指标上均达最优,在质量和效率上超越传统及多步生成方法,单视频去噪仅需约1秒。
英文摘要:
Video object removal is a fundamental yet challenging task in video editing. Despite recent progress, existing methods typically fall into two categories. Traditional approaches based on optical flow or attention mechanisms often introduce noticeable artifacts and yield unnatural results. In contrast, diffusion-based methods improve visual realism but demand multiple denoising steps, limiting their practicality. To address these issues, we propose From-Draft-to-Draft-Free (D2DF), a framework that distills the ability of transforming coarse drafts into refined videos into a one-step video generation model. Within D2DF, a teacher model is trained to refine low-quality removal results ("drafts") into high-fidelity videos by multiple steps. Then, through Prior-Privileged Consistency Distillation (PPCD), we distill this capability into a student model that performs one-step removal conditioned on the draft. To eliminate draft dependency, we introduce a Self-Guided Fast Planting (SGFP) module based on our Temporal Masked Transformer that autonomously generates scene-consistent pseudo-drafts in latent space, enabling a fully draft-free one-step model. Extensive experiments show that both draft-conditioned and draft-free versions achieve state-of-the-art performance on multiple metrics, surpassing traditional and multi-step generative methods in both quality and efficiency. The denoising process for a single video takes only about 1 second.