发表机构
Institute of Information Science, Beijing Jiaotong University; City University of Macau; Meitu Inc(北京交通大学信息科学研究所; 澳门城市大学; 美图公司)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对基于草图的图像编辑缺乏像素级精度及对应数据集、评估指标的问题,提出SI-Edit框架,构建SI-Data数据集并建立评估指标,实现更可靠的结构控制与像素级局部编辑效果。
AI 中文摘要
尽管生成式模型已取得快速进展,但在基于草图的图像编辑中实现像素级精度仍是持续存在的挑战,尤其是对于细粒度的局部变形。这一差距主要源于高质量公开基准数据集的严重匮乏,这类数据集需同时提供几何约束和语义指令。为解决该问题,我们首先推出**SI-Data**,这是一个专门为指令引导的局部草图编辑设计的高质量数据集。我们开发了一个利用多模态大语言模型(MLLMs)的自动化流程,用于合成包含原始图像、局部几何草图、语义指令及对应编辑后图像的完整四元组。通过提供可靠的空间锚点和明确的语义意图,SI-Data独特地支持空间-语义协同学习。在此基础上,我们提出了名为**SI-Edit**的协同框架,该框架将语义指令与精确的几何约束相集成。此外,为解决缺乏标准化评估的问题,我们建立了一套全面的指标,用于衡量结构保真度(例如草图到边缘的对齐)和语义契合度。实验结果表明,SI-Edit在基于草图的图像编辑中比基线方法提供更可靠的结构控制,并实现与用户意图一致的精确像素级局部优化。相关数据和代码已在项目页面(this https URL)发布。
英文摘要
Despite rapid advances in generative models, achieving pixel-level precision in sketch-based image editing remains a persistent challenge, particularly for fine-grained local deformations. This gap stems primarily from the critical shortage of high-quality, publicly available benchmark datasets that jointly provide geometric constraints and semantic instructions. To address this issue, we first introduce **SI-Data**, a high-quality dataset specifically designed for instruction-guided local sketch editing. We develop an automated pipeline leveraging Multimodal Large Language Models (MLLMs) to synthesize comprehensive quadruplets comprising original images, local geometric sketches, semantic instructions, and corresponding edited images. By providing both reliable spatial anchors and explicit semantic intent, SI-Data uniquely enables collaborative spatial-semantic learning. Building upon this, we propose a collaborative framework called **SI-Edit** that integrates semantic instructions with precise geometric constraints. Furthermore, to address the lack of standardized evaluation, we establish a comprehensive set of metrics designed to measure both structural fidelity (e.g., sketch-to-edge alignment) and semantic adherence. Experimental results demonstrate that SI-Edit provides more reliable structural control than baselines for sketch-based image editing, and achieves precise, pixel-level local refinements aligned with user intent. The data and code are released on the [project page](https://github.com/ywxsuperstar/SIEdit).
Commentsaccepted by ACM MM 2026