AI 中文总结
综述基于指令的图像编辑研究,围绕任务定义、数据构建、模型架构、评估指标及商业应用五个维度展开,提出综合深度诊断基准,通过开源方案比较揭示优缺点,探讨未来研究方向以推动该领域发展。
AI 中文摘要
基于指令的图像编辑(IIE)旨在依据文本指令将给定图像转化为新图像。大语言模型(LLMs)和视觉语言模型(VLMs)的进展加速了实用“单句图像编辑”系统的发展。本综述围绕五个核心维度对IIE研究进行系统分类和全面回顾,包括任务定义与编辑操作分层分类、训练数据构建方法、从基于GAN到扩散和自回归范式的架构演变、标准化评估指标与基准开发以及商业解决方案介绍。分析展示了各模型代的关键技术里程碑,还提出了IIE任务综合深度诊断基准(CDD-IIE Bench),通过开源解决方案的实证比较突出其各自优缺点,最后讨论了该领域未来研究方向。
英文摘要
Instruction-based Image Editing (IIE) aims to transform a given image into a new one based on textual instructions. Advances in Large Language Models (LLMs) and Vision-Language Models (VLMs) have accelerated progress toward practical ``one-sentence image editing" systems. This survey presents a systematic taxonomy and comprehensive review of IIE research, structured around five core dimensions: (1) task definition and hierarchical categorization of editing operations, (2) methodologies for training data construction, (3) architectural evolution from GAN-based to diffusion and autoregressive paradigms, (4) standardized evaluation metrics and benchmark development, and (5) introduction of commercial solutions. Our analysis shows critical technological milestones across model generations. We further propose a Comprehensive, in-Depth, and Diagnostic benchmark for IIE task (CDD-IIE Bench), which can rigorously assess the multiple aspects of model performance. Through empirical comparisons of open-source solutions, we highlight their respective capabilities and limitations. Finally, we discuss future research directions to advance the field.
Comments33 pages, 7 figures, Vicinagearth
Journal refZang, X., Jiang, Z., Cheng, J. et al. Instruction-based image editing: a survey on data, models, evaluation, and applications. Vicinagearth 3, 3 (2026)
DOI:10.1007/s44336-026-00034-3