发表机构
Zhejiang University; University at Buffalo; University of California, Berkeley(浙江大学; 纽约州立大学布法罗分校; 加州大学伯克利分校)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出DecFlowEdit,通过解耦定位与编辑的引导尺度,利用无CFG的速度差异先验重加权更新,实现无需训练和反转的基于流图像编辑,显著减少背景泄漏并保持编辑质量。
AI 中文摘要
基于流的图像编辑(FlowEdit)能够通过源速度与目标速度之间的差异实现无反转的语义变化。在本文中,我们观察到FlowEdit默认的分类器自由引导(CFG)配置,即源与目标尺度不对称,会导致严重的背景泄漏。匹配这些引导尺度,例如移除CFG,可以改善编辑相关的定位,但会严重降低可编辑性。为了兼顾两者,我们提出了DecFlowEdit,它在基于流的生成模型中解耦了用于定位和编辑的最优引导尺度。具体而言,DecFlowEdit首先通过时间聚合在无CFG条件下评估的速度差异来提取编辑相关先验,然后利用该先验在默认CFG下重新加权原始更新。我们的方法保持无需训练且无需反转,既不需要外部空间掩码,也不需要注意力操作。在PIE-Bench上针对FLUX、SD3和SD3.5的实验表明,DecFlowEdit改善了背景保持,相对于FlowEdit在可比较的编辑保真度下,结构距离减少了约61%至73%,背景LPIPS减少了68%至80%。
英文摘要
Flow-based image editing (FlowEdit) enables inversion-free semantic changes through the difference between source and target velocities. In this paper, we observe that FlowEdit's default classifier-free guidance (CFG) configuration, with asymmetric source and target scales, causes substantial background leakage. Matching these guidance scales, for example by removing CFG, improves edit-relevant localization but severely degrades editability. To get the best of both worlds, we propose DecFlowEdit, which decouples the optimal guidance scales for localization and for editing in flow-based generative models. In particular, DecFlowEdit first extracts an edit-relevant prior by temporally aggregating velocity differences evaluated without CFG, and then uses this prior to reweight the original updates under default CFG. Our method remains training-free and inversion-free, requiring neither external spatial masks nor attention manipulation. Experiments on PIE-Bench across FLUX, SD3, and SD3.5 show that DecFlowEdit improves background preservation, reducing structure distance by approximately 61 to 73 percent and background LPIPS by 68 to 80 percent relative to FlowEdit at comparable editing fidelity.