AI 中文总结
该研究针对视觉自回归模型(VARs)编辑的现有不足,提出以源为中心视角的EditMod方法,实现了高保真度与强文本对齐的1K图像端到端编辑,速度达1.57秒/张A100 GPU。
AI 中文摘要
下一代视觉自回归模型(VARs)已成为一种强大的生成范式,通过高效的 coarse-to-fine(由粗到细)预测生成高质量图像。然而,其在文本引导图像编辑方面的潜力仍未得到充分探索。现有的无训练VAR编辑方法通常将编辑表述为以目标为条件的再生,受源图像引导或约束,可能依赖于反演、测试时优化、注意力控制或用户提供的掩码。这种以生成为中心的表述未能充分利用VARs提供的多尺度源表示,且可能引入额外计算或干预。我们转而采用以源为中心的VAR编辑视角,其中编码的源图像标记作为主要视觉状态,编辑过程聚焦于条件诱导的变化。基于此视角,我们提出EditMod,该方法在共享自回归上下文下比较源条件和目标条件的预测,将二者的差异视为逐尺度编辑方向,并将其作为残差更新应用于选定尺度的源标记。实验表明,EditMod在保持强文本对齐的同时实现了领先的源图像保真度,且在单张A100 GPU上仅需1.57秒即可完成1K图像的端到端编辑,无需每张图像的准备工作。
英文摘要
Next-scale visual autoregressive models (VARs) have emerged as a powerful generative paradigm, producing high-quality images through efficient coarse-to-fine prediction. However, their potential for text-guided image editing remains largely underexplored. Existing training-free VAR editing approaches often formulate editing as target-conditioned regeneration guided or constrained by the source image, and may rely on inversion, test-time optimization, attention control, or user-provided masks. This generation-centric formulation does not fully exploit the multiscale source representations provided by VARs and may introduce additional computation or intervention. We instead take a source-centric perspective on VAR editing, in which the encoded source image tokens serve as the primary visual state and the editing process focuses on condition-induced changes. Based on this perspective, we propose \textbf{EditMod}, which compares source- and target-conditioned predictions under a shared autoregressive context, treats their difference as a scale-wise editing direction, and applies it as a residual update to source tokens at selected scales. Experiments show that EditMod achieves leading source-image fidelity while maintaining strong text alignment, and completes end-to-end editing of a 1K image in only 1.57 seconds on a single A100 GPU without per-image preparation.