发表机构
Hong Kong University of Science and Technology(香港科技大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
SoftPaint是一种零样本扩散模型编辑方法,利用软掩码实现像素级可调强度的图像与视频编辑,无需梯度且内存高效。
AI 中文摘要
基于提示词和参考图像引导的扩散模型编辑虽已取得快速进展,但在像素级控制上仍显粗糙。一个有前景的方向是引入软掩码来指定空间变化的编辑强度,但训练这种细粒度控制需要昂贵的像素级标注,而现有的零样本方法往往产生不尽人意的结果。我们提出了SoftPaint,一种新的零样本采样方法,利用软掩码实现从完全保留原始内容到完全重新合成掩码区域的连续编辑谱系。超越零样本修复方法,我们设计了一种基于Langevin迭代的采样器,尊重每个像素的软掩码强度,该方法普遍适用于图像和视频扩散模型,支持视频编辑等任务。该方法无需梯度,内存高效,并在多种图像和视频骨干网络上实现平滑的像素级编辑。
英文摘要
Diffusion models with prompt and reference image-guided editing have seen rapid progress, yet they remain too coarse for pixel-level control. One promising direction is to incorporate a soft mask that specifies spatially varying edit strengths but training such fine-grained control demands expensive pixel-wise annotations, while existing zero-shot methods often yield unsatisfactory results. We introduce SoftPaint, a new zero-shot sampling method that leverages soft masks to enable a continuous spectrum of edits, from fully preserving the original content to completely re-synthesizing the masked region. Going beyond zero-shot inpainting methods, we design a Langevin-iteration-based sampler that respects per-pixel soft mask strengths, which applies universally to image and video diffusion models, enabling tasks such as video editing. The method is gradient-free, memory-efficient, and achieves smooth, pixel-level edits across multiple image and video backbones.