细化本质上是可编辑的:基于生成式细化网络的无训练提示到提示图像编辑
Refinement Is Inherently Editable: Training-Free Prompt-to-Prompt Image Editing with Generative Refinement Network
查看机构详情
- City University of Hong Kong (Dongguan)(香港城市大学(东莞))
- City University of Hong Kong(香港城市大学)
- Mohamed bin Zayed University of Artificial Intelligence(穆罕默德·本·扎耶德人工智能大学)
机构由 AI 辅助整理,请以论文原文为准。
浏览论文内容
中文总结 AI 辅助
RefineEdit提出一种无训练的图像编辑框架,通过全局细化二进制码耦合编辑定位与内容生成,在PIE-Bench上实现最佳背景保留和CLIP分数。
中文摘要 AI 辅助
文本引导的图像编辑必须在引入所需更改的同时保留无关的源内容。基于扩散的编辑器依赖于空间控制,其不准确性可能导致编辑不完整或改变无关区域。因果自回归编辑器面临进一步的限制:其固定的解码顺序限制了对早期决策的修订。我们引入了RefineEdit,一个基于生成式细化网络的无训练提示到提示图像编辑框架。我们的关键思想是将编辑定位与内容生成通过二进制图像码的全局细化相结合,允许随着图像的演变重新评估编辑证据。RefineEdit从中间源状态初始化编辑分支,重用出现的布局。我们比较两个分支分配给相同源采样位的概率,使用它们的符号差异来选择可编辑的位置和位。选定的位遵循编辑细化,而剩余的位复制演变的源状态。为了在细化步骤中稳定编辑,自适应空间冻结限制了不必要的掩码扩展,而有限位锁定使最近选定的位保持可编辑。该框架不需要额外的训练、外部掩码或注意力控制。在PIE-Bench的九个编辑类别中,RefineEdit在PSNR、LPIPS、MSE和SSIM上取得了最佳的背景保留分数,以及在评估方法中取得了最高的整图和编辑区域CLIP分数。
英文摘要
Text-guided image editing must introduce the requested changes while preserving unrelated source content. In training-free editing, diffusion editors often use spatial controls whose inaccuracies can leave edits incomplete or alter unrelated regions. Causal autoregressive editors face a further constraint: their fixed decoding order limits revision of earlier decisions. As the first to explore training-free image editing with Generative Refinement Networks (GRN), we observe that its refinement process is inherently suitable for editing and offers a promising way to address these limitations. Motivated by this observation, we introduce RefineEdit, a training-free prompt-to-prompt image editing framework built on the GRN. Our key idea is to couple edit localization with content generation through the global refinement of binary image codes, allowing editing evidence to be revised as the image evolves. More specifically, RefineEdit combines bit routing with two stabilization mechanisms: adaptive spatial freezing and finite bit locking. Bit routing starts from an intermediate source state and uses signed probability differences between the two branches to identify editable positions and bits. It directs selected bits toward editing refinement while anchoring the rest to the evolving source trajectory. Adaptive spatial freezing limits unnecessary expansion of the editing region, while finite bit locking maintains recent bit activations to support continued editing. The overall framework requires no additional training, external masks, or attention control. Across nine editing categories of PIE-Bench, RefineEdit achieves the best background-preservation scores in PSNR, LPIPS, MSE, and SSIM, together with the highest whole-image and edited-region CLIP scores among the evaluated methods. Code is available at https://github.com/mura1n/RefineEdit.