arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.13861cs.CV

XSA-MAD:用于变形攻击检测的跨模态语义对齐

XSA-MAD: Cross-modal Semantic Alignment for Morphing Attack Detection

  • RIKEN AIP(理化学研究所人工智能项目)

机构由 AI 辅助整理,请以论文原文为准。

Jie Jin, Mahiro Tokumasu, Yu Makino, Masakatsu Nishigaki, Tetsushi Ohki

AI总结:

针对现有图像型变形攻击检测方法泛化性差的问题,提出基于CLIP的多模态框架XSA-MAD,通过对齐图像与语义空间实现对不同变形攻击的检测,在高保真攻击下性能优于现有方法

AI中文摘要:

变形攻击对人脸识别系统构成严重威胁。然而,现有的基于图像的变形攻击检测(MAD)方法仅依赖视觉线索,对未见过的生成技术泛化能力较差。我们提出XSA-MAD,一种基于CLIP的多模态框架,明确建模真实人脸与变形人脸之间的语义不一致性。变形概念被分解为四个可解释属性,包括身份、面部几何、纹理和一致性,并被编码为结构化且感知属性的文本表示。图像编码器与该判别性文本空间逐步对齐,得到统一的语义表示,捕捉真实图像与变形图像之间的生成不变性和概念级差异。在SMDD上训练后,于MAD22和MorDIFF上进行的实验表明,其对不同变形原理具有强泛化能力。特别是,XSA-MAD在基于GAN的变形攻击上达到2.92%的等错误率,且在高保真生成攻击下始终优于现有方法。

英文摘要:

Morphing attacks pose a serious threat to face recognition systems. However, existing image-based morphing attack detection (MAD) methods often generalize poorly to unseen generation techniques because they rely solely on visual cues. We propose XSA-MAD, a CLIP-based multimodal framework that explicitly models semantic inconsistencies between bona-fide and morphed faces. Morphing concepts are decomposed into four interpretable attributes, including identity, facial geometry, texture, and consistency, and are encoded as structured and attribute-aware textual representations. The image encoder is progressively aligned with this discriminative textual space, resulting in a unified semantic representation that captures generation-invariant and concept-level discrepancies between bona-fide and morph images. Experiments on MAD22 and MorDIFF, following training on SMDD, demonstrate strong generalization across diverse morphing principles. In particular, XSA-MAD achieves an equal error rate of 2.92% on GAN-based morphs and consistently outperforms existing methods under high-fidelity generative attacks.

补充信息

↑