arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

RIDGE:用于图像编辑的基于内部动态引导的重加噪方法

RIDGE: Re-Noising with Internal Dynamic Guidance for Image Editing

Ruiliang Gong, Zhen Wang, Yanghao Wang, Long Chen

arXiv 2608.03059首次发表:更新:

AI 中文总结

该研究针对无反转图像编辑中等位移构造的缺陷,提出RIDGE方法,通过重加噪与内部动态引导优化编辑,在多基准实验中实现了源保留、目标对齐与感知质量的良好权衡。

AI 中文摘要

无反转的基于流的图像编辑避免了潜在反转,但仍需在每一步编辑中获取目标侧状态。广泛使用的等位移构造方法,使噪声源状态与目标侧状态间的位移在不同噪声水平下保持不变,这与加噪过程不一致——在加噪时,用相同噪声水平和噪声样本加噪的两个干净状态间的位移应随噪声水平升高而收缩,因此该方法会在高噪声水平下导致过度激进的更新。我们提出RIDGE:用于图像编辑的基于内部动态引导的重加噪方法,这是一种无反转、无需训练的方法,它将编辑状态维持为不可用干净目标状态的演化近似。RIDGE使用与干净源状态相同的噪声水平和噪声样本对该近似进行重加噪,使它们的噪声位移随噪声升高自然减小。由于编辑状态初始包含的目标语义有限,RIDGE会在早期高噪声步骤中进一步应用内部动态引导:通过模型内部导出的软动态掩码,用干净目标状态预测结果引导临时编辑状态,将引导聚焦于需要修改的区域,无需外部分割或检测模型。在使用SD3 Medium和FLUX.1-dev两种骨干网络的两个基准上的实验表明,RIDGE在源内容保留、目标对齐和感知质量之间提供了良好的综合权衡。

英文摘要

Inversion-free flow-based image editing avoids latent inversion, but still requires a target-side state at every editing step. The widely used equal-displacement construction keeps the displacement between the noisy source state and the target-side state unchanged across noise levels. This is inconsistent with noising, under which the displacement between two clean states noised with the same noise level and noise sample should contract as the noise level increases. Thus, it can lead to overly aggressive updates at high noise levels. We introduce RIDGE: Re-Noising with Internal Dynamic Guidance for Image Editing, an inversion-free and training-free method that maintains the edited state as an evolving approximation to the unavailable clean target state. RIDGE re-noises this approximation using the same noise level and noise sample as the clean source state, allowing their noisy displacement to decrease naturally with increasing noise. Since the edited state initially contains limited target semantics, RIDGE further applies internal dynamic guidance during the early high-noise steps. A clean target state prediction guides the provisional edited state through a soft dynamic mask derived internally from the model, focusing guidance on regions that require modification without external segmentation or detection models. Experiments on two benchmarks using two backbones, SD3 Medium and FLUX.1-dev, show that RIDGE offers a favorable aggregate trade-off among source preservation, target alignment, and perceptual quality.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑