AI 中文总结
研究针对现有一步图像编辑方法缺乏空间控制及局部编辑能力不足的问题,提出WhereEdit框架,通过自动识别语义相关区域并应用自适应局部调制,实现局部潜在编辑,实验证明其优于现有方法,提升了一步图像编辑质量。
AI 中文摘要
近期的一步文本到图像(T2I)模型实现了高效图像合成,并为实时图像编辑带来新机遇。然而,现有一步编辑方法主要依赖文本条件进行语义转换,缺乏对编辑位置的明确空间控制。更重要的是,即便引入空间约束,这些方法在目标区域内也难以实现强大且稳定的语义修改。本文从空间控制角度重新审视一步图像编辑,识别出两个关键挑战:发现可编辑区域和实现有效的局部语义转换。我们揭示现有方法进行全局语义传输,限制了一步设置下的高强度局部编辑。为解决此问题,我们提出WhereEdit框架,将一步编辑重新定义为局部自适应编辑。WhereEdit自动从内部模型特征中识别语义相关区域,并应用自适应局部调制来增强目标区域编辑,同时保留非目标区域和结构一致性。在PIE - Bench基准上的实验表明,WhereEdit始终优于现有一步图像编辑方法,在保持一步生成效率的同时实现了更高的编辑质量。区域级监督的额外实验进一步凸显了明确空间推理对高质量一步图像编辑的重要性。
英文摘要
Recent one-step text-to-image (T2I) models enable efficient image synthesis and provide new opportunities for real-time image editing. However, existing one-step editing methods primarily rely on text conditioning for semantic transformation, lacking explicit spatial control over \textit{where} to edit. More importantly, even when spatial constraints are introduced, these methods often struggle to achieve strong and stable semantic modifications within the target regions. In this work, we revisit one-step image editing from a spatially controlled perspective and identify two key challenges: discovering editable regions and achieving effective localized semantic transformation. We reveal that existing methods perform global semantic transport, which limits high-intensity local editing under the one-step setting. To address this issue, we propose \textbf{WhereEdit}, a framework that reformulates one-step editing as localized adaptive editing. WhereEdit automatically identifies semantically relevant regions from internal model features and applies adaptive local modulation to enhance target-region editing while preserving non-target areas and structural consistency. Experiments on the PIE-Bench benchmark demonstrate that WhereEdit consistently outperforms existing one-step image editing methods, achieving superior editing quality while maintaining the efficiency of one-step generation. Additional experiments with region-level supervision further highlight the importance of explicit spatial reasoning for high-quality one-step image editing.