发表机构
School of Engineering, Universidad Autónoma de Madrid(马德里自治大学工程学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对真实ID数字篡改检测的领域差距,构建首个含520万补丁的FakeIDet3-DB数据库,提出PACE算法,评估显示现有模型对其攻击检测定位性能不佳。
AI 中文摘要
身份文档(ID)认证依赖于复杂高频安全图案的结构完整性。然而,先进的生成式AI模型如今可注入局部高保真篡改,形成欺骗性攻击以绕过标准验证。训练鲁棒的图像取证模型检测这些异常受隐私法规阻碍,迫使依赖缺乏真实ID复杂视觉图案的合成模板。为弥合该领域差距,我们推出FakeIDet3-DB,首个针对真实政府签发ID的数字篡改综合数据库。FakeIDet3-DB涵盖经典篡改(如复制-移动)与生成式AI驱动篡改(如换脸、图像修复),并采用高级图像细化流程抑制视觉伪影。此外,为符合GDPR等严格数据保护法规,我们采用近期提出的基于补丁的框架。为最大化取证效用,我们将从真实ID提取隐私感知补丁建模为几何约束图像处理问题,提出PACE(Pseudo-Anonymized Contextual patch Extraction算法),其利用积分图像映射与距离驱动非极大值抑制(NMS),高效勾勒匿名化掩码以防止个人身份信息(PII)泄露,同时最大化审查周边区域的语义密度,最终从超6400张真实/伪造ID图像中提取近520万个补丁。进一步,我们采用最先进模型对FakeIDet3-DB开展广泛评估,结果显示这些模型均难以检测和定位来自生成式与经典技术的攻击,检测等错误率(EER)为32.45%,定位的受试者工作特征曲线下面积(AUC-ROC)为83.48%。
英文摘要
Identity document (ID) authentication relies on the structural integrity of complex, high-frequency security patterns. However, advanced Generative AI models can now inject localized, high-fidelity manipulations, creating deceptive attacks that bypass standard verification. Training robust image forensic models to detect these anomalies is hindered by privacy regulations, forcing reliance on synthetic templates lacking the intricate visual patterns of real IDs. To bridge this domain gap, we introduce FakeIDet3-DB, the first comprehensive database of digital manipulations on real, government-issued IDs. FakeIDet3-DB encompasses classical (e.g., copy-move) and Generative AI-driven manipulations (e.g., face-swapping, inpainting) enhanced with advanced image refinement procedures to suppress visual artifacts. In addition, to comply with strict data protection regulations (e.g., GDPR), we adopt a recently-proposed framework based on patches. In order to maximize forensic utility, we formulate privacy-aware patch extraction from a real ID as a geometrically constrained image processing problem. We propose PACE, a Pseudo-Anonymized Contextual patch Extraction algorithm, which leverages Integral Image mapping and distance-driven Non-Maximum Suppression (NMS). PACE efficiently contours anonymization masks that prevent Personally Identifiable Information (PII) leakage while maximizing semantic density in peri-censorship regions, yielding almost 5.2M patches extracted from more than 6.4K images from real/fake IDs. Furthermore, an extensive evaluation of the proposed FakeIDet3-DB is performed using state-of-the-art models, showcasing they all struggle to detect and locate attacks coming from generative and classic techniques (32.45\% EER in detection and 83.48\% AUC-ROC in localization).
Comments13 Pages, Database, Pipeline, Algorithm