发表机构
KAUST; Krea AI(阿卜杜拉国王科技大学; Krea AI)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对多轮图像编辑中模型因条件分布不匹配而性能退化的问题,提出在线策略自蒸馏框架MT-OPSD,利用干净条件教师监督训练自生成状态,并引入LME-Bench基准,实验证明其显著提升长时程编辑成功率并减少崩溃。
AI 中文摘要
基于指令的图像编辑在单轮设置中已取得强劲性能,然而实际编辑往往是迭代式的,每条指令都应用于上一轮的输出。我们发现现有编辑模型在递归编辑下性能迅速退化,并将这一失败归因于条件分布中的训练-测试不匹配:模型在干净的源图像上训练,但在推理时必须反复以其自身不完美的输出为条件。为解决此问题,我们提出MT-OPSD,一种在线策略自蒸馏框架,该框架使用来自干净条件教师模型的编辑监督来训练模型处理自生成的条件状态,而无需多轮标注。我们进一步引入LME-Bench,一个包含100个十轮编辑会话的基准,用于评估长时程鲁棒性。在三个编辑骨干网络上的实验表明,MT-OPSD显著提高了长时程编辑成功率,减少了多轮崩溃,同时基本保持了单轮编辑质量。
英文摘要
Instruction-based image editing has achieved strong performance in single-turn settings, yet practical editing is often iterative, with each instruction applied to the output of the previous turn. We find that existing editing models degrade rapidly under recursive editing and attribute this failure to a train-test mismatch in the conditioning distribution: models are trained on clean source images but must repeatedly condition on their own imperfect outputs at inference time. To address this, we propose MT-OPSD, an on-policy self-distillation framework that trains the model on self-generated conditioning states with editing supervision from a clean-conditioned teacher, without requiring multi-turn annotations. We further introduce LME-Bench, a benchmark of 100 ten-turn editing sessions for evaluating long-horizon robustness. Experiments across three editing backbones show that MT-OPSD substantially improves long-horizon editing success and reduces multi-turn collapse while largely preserving single-turn editing quality.