发表机构
Guangdong University of Technology; Hong Kong Baptist University; Peking University; Huizhou University(广东工业大学; 香港浸会大学; 北京大学; 惠州学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对文本引导扩散图像编辑的全局漂移问题,提出推理时框架ATDEdit,通过逐令牌条件意外度估计可编辑位置,在PIE-Bench上实现最优保留指标且语义对齐性具竞争力。
AI 中文摘要
文本引导的扩散图像编辑旨在修改图像的语义属性,同时保留其身份、布局和背景。然而,在采样过程中直接切换文本条件常会导致全局漂移,因为去噪动力学的变化会在令牌间传播,可能破坏未编辑区域。为解决该问题,我们提出异步令牌解码编辑(ATDEdit),这是一种推理时框架,将每个采样器步骤视为全局耦合令牌矩阵的并行更新,支持按令牌索引切换条件并采用差异化更新策略。ATDEdit不对所有令牌应用同步的目标条件更新,而是通过逐令牌条件意外度估计可编辑位置,并对选中的令牌集应用目标条件修正;它在保留令牌位置提供源键/值记忆,并将选中的保留令牌潜在行投影回其源值,这些操作可促进背景保留,但不构成像素级不变性保证。该方法结合了局部编辑与背景保留,无需外部或用户提供的空间掩码,也无需模型微调。在PIE-Bench上,ATDEdit取得了报告中最优的保留指标,包括27.44 dB的PSNR和0.055的LPIPS,同时保持了具有竞争力的语义对齐效果。
英文摘要
Text-guided diffusion image editing aims to modify semantic attributes of an image while preserving its identity, layout, and background. However, naïvely switching the text condition during sampling often causes global drift, as denoising dynamics propagate changes across tokens and can disrupt unedited regions. To address this issue, we propose \textbf{A}synchronous \textbf{T}oken \textbf{D}ecoding \textbf{Edit} (ATDEdit), an inference-time framework that views each sampler step as a parallel update of a globally coupled token matrix and enables token-indexed condition switching with differentiated update policies. Instead of applying synchronous target-conditioned updates to all tokens, ATDEdit estimates editable locations using token-wise conditional surprisal and applies target-conditioned corrections to the selected token set. It supplies source key/value memory at keep-token positions and projects selected keep-token latent rows back to their source values; these operations promote background preservation but do not constitute a pixel-level invariance guarantee. This approach combines local editing and background preservation without external or user-provided spatial masks and without model fine-tuning. On PIE-Bench, ATDEdit achieves the strongest reported preservation metrics, including 27.44~dB PSNR and 0.055 LPIPS, while retaining competitive semantic alignment.
CommentsAccepted by ACMMM 2026