arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

FUSED:面向AI图像修复检测与定位的取证语义混合专家模型

FUSED: Forensic-Semantic Mixture-of-Experts for AI Inpainting Detection and Localization

Anton Nuzhdin, Marcel Worring, Ivona Najdenkoska

arXiv 2608.28302首次发表:更新:

发表机构

Informatics Institute, University of Amsterdam(阿姆斯特丹大学信息学研究所)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对AI图像修复检测与定位的分布偏移问题,提出FUSED框架,结合取证线索与语义特征,在跨生成器基准上实现最优检测与定位,迁移性能显著提升。

AI 中文摘要

基于扩散的图像修复模型仅修改图像的局部区域,而许多AI图像检测器依赖全局伪影且无法定位,这些伪影会因生成器不同而变化,限制了分布偏移下检测器的迁移能力。近期研究表明,修复图像修复区域外的真实像素可消除这些线索并降低预训练检测器的性能。为解决该问题,我们提出FUSED,这是一个用于联合检测和定位AI生成图像修复的统一框架。FUSED采用稀疏门控混合专家架构,将低级取证线索与高级语义特征相结合,使模型能为每个token自适应地优先选择最相关的信号。对于每个输入,FUSED同时预测图像级的篡改分数和修复区域的像素级掩码。在OpenSDID跨生成器基准测试中,FUSED取得了最佳的平均检测和定位性能,在未见过的生成器上的提升最为显著。同一模型直接迁移到预留的AutoSplice和CocoGlide基准测试中,定位性能提升超过一倍。对每个预留基准测试在有无全局生成器伪影的情况下进行评估,进一步表明所有被评估方法(包括我们的方法)都会部分将伪影解读为篡改的证据,而FUSED在两种条件下均保持最强性能。代码和预训练模型可在此https URL获取。

英文摘要

Diffusion-based inpainting models modify only a localized part of an image, while many AI-image detectors rely on global artifacts and do not localize. These artifacts vary across generators, limiting detector transfer under distribution shifts. Recent work shows that restoring the authentic pixels outside the inpainted region removes these cues and can degrade pretrained detectors. To address this, we present FUSED, a unified framework for the joint detection and localization of AI-generated inpainting. FUSED combines low-level forensic cues with high-level semantic features using a sparsely-gated Mixture-of-Experts architecture, enabling the model to adaptively prioritize the most relevant signal for each token. For each input, FUSED predicts both an image-level manipulation score and a pixel-level mask of the inpainted area. On the OpenSDID cross-generator benchmark, FUSED achieves the best average detection and localization, with the largest gains on unseen generators. The same model transfers directly to the held-out AutoSplice and CocoGlide benchmarks, more than doubling localization performance. Evaluating each held-out benchmark with and without the global generator artifact further shows that all evaluated methods, ours included, partly read the artifact as evidence of manipulation, and FUSED remains the strongest under both conditions. Code and pretrained models are available at https://github.com/AntonNuzhdin/FUSED.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑