SIGMA: Semantic-Difference Instruction-Grounding Mask Annotator for Text-Driven Image Manipulation Localization
SIGMA: 基于语义差异的指令引导掩码标注器用于文本驱动图像操作定位
机构 * Shenzhen Campus of Sun Yat-sen University(中山大学深圳校区) ; Guangdong Provincial Key Laboratory of Intelligent Information Processing and Shenzhen Key Laboratory of Media Security(广东省智能信息处理重点实验室和深圳媒体安全重点实验室) ; Shenzhen University of Advanced Technology and Shenzhen Institute of Advanced Technology, Chinese Academy of Sciences(深圳先进技术大学和深圳先进技术研究所,中国科学院) ; Alibaba Group(阿里巴巴集团) ; Shenzhen MSU-BIT University(深圳MSU-BIT大学)
专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV
AI总结 提出SIGMA方法,通过视觉基础模型中的语义特征差异和指令引导的空间先验,自动从公开编辑数据集中生成像素级掩码,用于训练图像操作定位模型,在五个基准上F1提升12.20%,并生成约110万训练集使六个检测器平均F1提升18.34%。