arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

建模编辑,而非图像:以源为中心视角的视觉自回归编辑

Model the Edit, Not the Image: Visual Autoregressive Editing from a Source-Centric Perspective

Hongyi Fang, Chuwen Xie, Benjia Zhou, Yu-Xuan Qiu, Chenggong Hu, Zhibin Wang, Chao Chen, Jianbin Qin, Rui Mao

arXiv 2608.09057首次发表:更新:

AI 中文总结

该研究针对视觉自回归模型(VARs)编辑的现有不足,提出以源为中心视角的EditMod方法,实现了高保真度与强文本对齐的1K图像端到端编辑,速度达1.57秒/张A100 GPU。

AI 中文摘要

下一代视觉自回归模型(VARs)已成为一种强大的生成范式,通过高效的 coarse-to-fine(由粗到细)预测生成高质量图像。然而,其在文本引导图像编辑方面的潜力仍未得到充分探索。现有的无训练VAR编辑方法通常将编辑表述为以目标为条件的再生,受源图像引导或约束,可能依赖于反演、测试时优化、注意力控制或用户提供的掩码。这种以生成为中心的表述未能充分利用VARs提供的多尺度源表示,且可能引入额外计算或干预。我们转而采用以源为中心的VAR编辑视角,其中编码的源图像标记作为主要视觉状态,编辑过程聚焦于条件诱导的变化。基于此视角,我们提出EditMod,该方法在共享自回归上下文下比较源条件和目标条件的预测,将二者的差异视为逐尺度编辑方向,并将其作为残差更新应用于选定尺度的源标记。实验表明,EditMod在保持强文本对齐的同时实现了领先的源图像保真度,且在单张A100 GPU上仅需1.57秒即可完成1K图像的端到端编辑,无需每张图像的准备工作。

英文摘要

Next-scale visual autoregressive models (VARs) have emerged as a powerful generative paradigm, producing high-quality images through efficient coarse-to-fine prediction. However, their potential for text-guided image editing remains largely underexplored. Existing training-free VAR editing approaches often formulate editing as target-conditioned regeneration guided or constrained by the source image, and may rely on inversion, test-time optimization, attention control, or user-provided masks. This generation-centric formulation does not fully exploit the multiscale source representations provided by VARs and may introduce additional computation or intervention. We instead take a source-centric perspective on VAR editing, in which the encoded source image tokens serve as the primary visual state and the editing process focuses on condition-induced changes. Based on this perspective, we propose \textbf{EditMod}, which compares source- and target-conditioned predictions under a shared autoregressive context, treats their difference as a scale-wise editing direction, and applies it as a residual update to source tokens at selected scales. Experiments show that EditMod achieves leading source-image fidelity while maintaining strong text alignment, and completes end-to-end editing of a 1K image in only 1.57 seconds on a single A100 GPU without per-image preparation.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑