发表机构
School of Artificial Intelligence, Shanghai Jiao Tong University; Qwen Business Unit of Alibaba; National University of Singapore(上海交通大学人工智能学院; 阿里巴巴通义千问业务部; 新加坡国立大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对现有扩散图像编辑模型无法高效处理超高清图像的问题,提出EditBridge扩散桥框架,通过先验引导的块级稀疏注意力机制实现4K高保真编辑,速度提升显著。
AI 中文摘要
专业工作流对高分辨率图像编辑的需求日益增长,但现有基于扩散的模型因二次注意力复杂度和过高的内存需求,仍受限于1K以下的分辨率。一种普遍的解决方案采用两阶段流水线:在低分辨率(LR)下编辑,随后进行独立的超分辨率处理。然而,该方法存在两个关键问题:信息发散,即生成的幻觉细节与原始高分辨率(HR)源图像矛盾;以及纹理退化,表现为过度平滑或过度锐化的伪影。我们提出EditBridge,一种用于高效超高清编辑的扩散桥框架。与传统从噪声中再生的扩散模型不同,我们将细化过程表述为从低分辨率编辑结果到其高分辨率对应结果的结构化数据到数据转换,明确以原始高分辨率源为条件,以保留真实细节。为高效融入高分辨率源引导,我们引入一种先验引导的块级稀疏注意力机制,该机制利用第一阶段编辑的语义对应关系,将跨图像交互约束在空间对齐区域,显著降低计算开销。大量实验表明,EditBridge在最高4K分辨率下实现了高保真编辑与卓越的感知质量,在2K分辨率下实现了3.6至8.4倍的加速,并能在61秒内完成实用的4K编辑。
英文摘要
High-resolution image editing is increasingly demanded in professional workflows, yet existing diffusion-based models remain constrained to resolutions below 1K due to quadratic attention complexity and prohibitive memory requirements. A prevalent workaround employs a two-stage pipeline: editing at low resolution followed by independent super-resolution. However, this approach suffers from two critical issues: information divergence, where hallucinated details contradict the original high-resolution (HR) source, and texture degradation, manifesting as over-smoothed or over-sharpened artifacts. We propose EditBridge, a diffusion bridge framework for efficient ultra high-resolution editing. Unlike conventional diffusion that regenerates from noise, we formulate refinement as structured data-to-data translation from the low-resolution (LR) edited result to its HR counterpart, explicitly conditioned on the original HR source to preserve authentic details. To efficiently incorporate HR source guidance, we introduce a prior-guided block-wise sparse attention mechanism that exploits semantic correspondence from first-stage editing to constrain cross-image interactions to spatially aligned regions, significantly reducing computational overhead. Extensive experiments demonstrate that EditBridge achieves high-fidelity editing with superior perceptual quality at resolutions up to 4K, delivering 3.6--8.4$\times$ speedup at 2K and enabling practical 4K editing in 61 seconds.