RegionCache:面向高效多轮图像生成的语义感知区域复用
RegionCache: Semantic-Aware Region Reuse for Efficient Multi-Turn Image Generation
浏览论文内容
中文总结 AI 辅助
针对多轮图像编辑中现有DiT方法的冗余计算问题,提出RegionCache框架,通过语义感知的区域复用实现1.43-2.55倍端到端加速且保持图像质量。
中文摘要 AI 辅助
现实场景中的图像生成常涉及多轮编辑,用户迭代修改小区域时,大部分图像内容保持不变。然而,现有基于扩散Transformer(DiT)的编辑流程每轮都会重新计算整个图像,造成大量冗余计算。现有DiT加速方法还忽略了不同提示间的语义对应关系,导致不必要的重新计算或不安全的复用,损害编辑质量。为解决该问题,我们提出RegionCache,一种面向多轮图像编辑的语义感知复用框架,可选择性复用未修改区域的扩散状态。RegionCache通过连续提示间的语义重叠和交叉注意力定位检测可复用区域,并基于提示相似度和上下文一致性采用自适应复用调度。在PixArt-alpha上的实验表明,RegionCache实现了1.43倍至2.55倍的端到端加速,同时保持相当的图像质量。代码可在指定URL获取。
英文摘要
Real-world image generation often involves multi-turn editing, where users iteratively modify small regions while most image content remains unchanged. However, existing diffusion transformer (DiT)-based editing pipelines recompute the entire image at every turn, causing substantial redundant computation. Existing DiT acceleration methods further ignore semantic correspondence across prompts, leading to unnecessary recomputation or unsafe reuse that harms editing quality. To address this, we propose RegionCache, a semantic-aware reuse framework for multi-turn image editing that selectively reuses diffusion states from unchanged regions. RegionCache detects reusable regions through semantic overlap between consecutive prompts and cross-attention localization, and adopts an adaptive reuse schedule based on prompt similarity and contextual consistency. Experiments on PixArt-alpha demonstrate that RegionCache achieves 1.43x--2.55x end-to-end speedup while maintaining comparable image quality. Code is available at https://github.com/hebutBryant/RegionCache.
发表机构
- School of Computer Science and Engineering, Northeastern University(东北大学计算机科学与工程学院)
机构由 AI 辅助整理,请以论文原文为准。