发表机构
HSE University; Yandex(高等经济大学; Yandex)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究提出任务切换方法,让统一编辑模型在去噪过程中利用文生图能力,在三个编辑器和四个基准上提升了编辑质量,同时保持感知保留度。
AI 中文摘要
统一模型同时针对基于指令的图像编辑和文生图(T2I)生成进行训练,但标准编辑流程在整个去噪过程中保持源图像的条件作用。我们探究编辑是否能从T2I中受益,并研究条件作用的效果如何随编辑类型和去噪阶段而变化。在纯编辑中,某些编辑的源注意力在采样轨迹上会下降。这一观察促使我们提出任务切换方法,使模型能够利用其T2I能力。在三个统一编辑器和四个基准测试中,在有限区间内切换到T2I任务可提高编辑质量,而三个模型的平均感知保留度仍接近纯编辑。因此,统一编辑器受益于使用其训练所得的两种条件模式,而切换时机则决定了质量与保留度之间的平衡。
英文摘要
Unified models are trained for both instruction-based image editing and text-to-image (T2I) generation, but standard editing pipelines keep source-image conditioning throughout denoising. We ask whether editing can benefit from T2I, and study how the effects of conditioning vary across edits and denoising stages. In pure editing, source attention declines for some edits over the sampling trajectory. This observation led us to task switching, which lets the model draw on its T2I capabilities. Across three unified editors and four benchmarks, switching to the T2I task for bounded intervals improves edit quality, while mean perceptual preservation remains close to pure editing across all three models. Unified editors therefore benefit from using both conditioning modes they are trained for, and the timing of the switch sets the balance between quality and preservation.
CommentsUnder review as a conference paper at ICLR 2027