arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

我们能显式地建模伪影吗?通过成对编辑关系解耦伪影用于图像篡改定位

Can We Model the Artifacts Explicitly? Disentangle Artifacts via Pairwise Edit Relations for Image Manipulation Localization

Xuekang Zhu, Kaiwen Feng, Ruifeng Wang, Xiwen Wang, Xiaochen Ma, Bo Du, Changjiang Jiang, Chenfan Qu, Songyu Ye, Xia Du, Wentao Feng, Jian Liu, Ji-Zhe Zhou

arXiv 2610.07916首次发表:更新:

发表机构

Sichuan University; Ant Group; The Hong Kong University of Science and Technology; Wuhan University; South China University of Technology; University of Southern California; Xiamen University of Technology(四川大学; 蚂蚁集团; 香港科技大学; 武汉大学; 华南理工大学; 南加州大学; 厦门理工学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对图像篡改定位中隐式伪影建模的不足,提出两阶段成对伪影学习范式,通过特征解纠缠显式建模伪影,并构建EditGroup-45K数据集,在多种架构上取得一致改进。

AI 中文摘要

图像篡改定位(IML)通常被表述为一个全监督学习任务,即估计给定图像 $x$ 的最优篡改掩码 $y$。在这项工作中,我们首先揭示伪影的潜在性质,从而将IML重新解释为一个潜变量问题,$P(y|x)=\int P(y|z)\\,P(z|x)\\,dz$,其中 $z$ 表示伪影。根据这一解释,我们指出当前IML模型不足的原因在于其隐式的伪影建模策略,强调以显式方式建模 $z$ 的必要性。在没有直接标签的情况下,特征解纠缠是这种显式建模最合适的解决方案。因此,我们提出了一种两阶段学习范式,包括成对伪影学习(PAL)和标准定位(SL)阶段,通过编辑关系估计 $P(z|x)$ 和 $P(y|z)$。为支持基于编辑关系的学习,我们进一步整理了EditGroup-45K,一个以源图像为锚点的数据集,按编辑组组织以构建图像对。大量实验表明,我们的PAL范式在多种IML架构上带来了一致的改进,实证分析进一步验证了PAL确实通过特征解纠缠显式捕获了伪影。代码和数据集可在以下网址获取:https://this URL

英文摘要

Image Manipulation Localization (IML) is commonly formulated as a fully supervised learning task that estimates the optimal manipulation mask $y$ for a given image $x$. In this work, we first reveal the latent nature of artifacts and thus reinterpret IML as a latent-variable problem, $P(y|x)=\int P(y|z)\,P(z|x)\,dz$, where $z$ denotes the artifacts. Following this interpretation, we pinpoint the cause for the current IML models' insufficiency as their implicit artifacts modeling strategy, highlighting the necessity of modeling $z$ in an explicit manner. Without direct labels, feature disentanglement is the most appropriate solution for this explicit modeling. Accordingly, we propose a two-stage learning paradigm with the Pairwise Artifacts Learning (PAL) and Standard Localization (SL) phases to estimate $P(z|x)$ and $P(y|z)$ via edit relations. To support our edit-relation-based learning, we further curate EditGroup-45K, a source-anchored dataset organized into edit groups for pair construction. Extensive experiments show that our PAL paradigm yields consistent improvements across diverse IML architectures, and empirical analyses further verify that PAL does capture artifacts explicitly through feature disentanglement. Code and dataset are available at https://github.com/venus-guangjian/PAL

CommentsNeurIPS 2026 (Oral)

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑