arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

通过概念缩放与密集监督释放图像编辑的潜力

Unlocking the Potential of Image Editing via Concept Scaling and Dense Supervision

Long Cui, Xiaoqian Liu, Qi Qin, Yi Xin, Tao Lin, Jianguo Li, Linfeng Zhang

arXiv 2608.16812首次发表:更新:

AI 中文总结

针对图像编辑框架存在的概念粒度关注不足与训练效率低的问题,构建了含1200万编辑对的ConceptEdit-12M数据集,提出密集监督训练策略,推出细粒度评估套件ConceptEdit-Bench,性能优于现有工作。

AI 中文摘要

现有图像编辑框架大多遵循文本到图像扩散模型的训练范式,但将该范式扩展到图像编辑时,会凸显两个固有差异:一是对编辑概念粒度的关注不足,二是由稀疏监督信号导致的训练效率低下。为解决这些问题,我们建立了包含1000多个细粒度编辑概念的综合分层分类体系,并通过改进的合成框架构建了ConceptEdit-12M——一个拥有1200万高质量编辑对的大型数据集。这种基于库的方法有效纠正了生成数据的分布崩溃,同时确保了高数据保真度。此外,我们提出了一种密集监督训练策略,该策略将多个互不干扰的概念合成到单个图像对中,通过提供更丰富的学习信号,显著提升了训练效率和整体模型性能。训练结果验证了我们的策略,其表现显著优于现有工作。最后,我们推出了ConceptEdit-Bench,这是一个用于诊断模型在大量真实场景下能力的细粒度评估套件。

英文摘要

Existing image editing frameworks predominantly follow the training paradigm of text-to-image diffusion models. However, extending this paradigm to image editing highlights two inherent discrepancies, specifically, the insufficient attention to edit concept granularity and the training inefficiency caused by sparse supervision signals. To address these issues, we establish a comprehensive hierarchical taxonomy featuring over 1,000 fine-grained edit concepts and build ConceptEdit-12M, a massive dataset of 12 million high-quality editing pairs via an improved synthesis framework. This library-driven approach effectively rectifies the distribution collapse of generated data while ensuring high data fidelity. Furthermore, we propose a dense supervision training strategy that synthesizes multiple non-interfering concepts into single image pairs. By providing richer learning signals, this strategy significantly enhances both training efficiency and overall model performance. Training results validate our strategy, significantly outperforming prior works. Finally, we present ConceptEdit-Bench, a granular evaluation suite designed to diagnose model capabilities across a vast array of real-world scenarios.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑