发表机构
Faculty of Engineering and Computer Science, Qaiwan International University; Department of Information and Communication Technology Center (ICTC)-System Information, Ministry of Higher Education and Scientific Research(工程与计算机科学学院,卡伊万国际大学; 信息与通信技术中心(ICTC)-系统信息,高等教育与科学研究部)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对传统人脸识别系统在面部遮挡时的局限,提出PLGSA-Transformer框架。利用眼周地标引导注意力、混合CNN-Transformer分支建模及遮挡自适应余弦阈值,提升跨模态蒙面人脸识别性能。
AI 中文摘要
COVID-19加速了面部口罩的广泛使用,这暴露了传统人脸识别系统的局限性。现有方法在面部区域被遮挡时无法泛化。本文提出PLGSA-Transformer,它有三点贡献:眼周地标引导空间注意力、混合CNN-Transformer分支和遮挡自适应余弦阈值。该模型在多数据集上评估,性能超越其他方法。
英文摘要
The widespread adoption of facial masks, accelerated by COVID-19 and mandated in security-sensitive settings, has exposed limitations of conventional face recognition systems. Existing approaches relying on fixed cosine thresholds, non-adaptive CNNs, and purely data-driven features fail to generalize when facial regions are occluded, creating a gap between lab performance and real-world deployability. This paper proposes PLGSA-Transformer, a cross-modal face matching framework with three contributions. First, Periocular Landmark-Guided Spatial Attention (PLGSA) uses MediaPipe landmarks to compute Gaussian heatmaps over the eye, brow, and forehead regions, fusing them with EfficientNetB3 features via a learnable residual gate to direct attention toward discriminative visible regions. Second, a Hybrid CNN-Transformer Branch reshapes feature maps into tokens processed by a two-layer Multi-Head Self-Attention encoder, enabling cross-regional dependency modelling. Third, the Occlusion-Adaptive Cosine Threshold (OACT) is a jointly trained head that raises the matching threshold in proportion to predicted occlusion severity. The model is evaluated on 858 images from Zenodo MDMFR (60%), Kaggle CelebA-HQ masked collection (25%), and author-collected images (15%), spanning both genders, ages 21-75, with varied mask types, trained via a unified loss combining contrastive verification, identity classification, and occlusion cross-entropy. PLGSA-Transformer achieves 97.22% pair verification accuracy with ROC AUC 1.0000, surpassing VGG-16-based MUFM (Abdullah et al., 2025; 95.0%), HOG classifiers (Adnan et al., 2020; 85.0%), and Feature-based Structural Measure (Shnain et al., 2017; 86.61%). These results confirm that encoding periocular geometry into attention, with Transformer modelling and occlusion-adaptive thresholds, yields a robust, scalable solution for cross-modal masked face recognition.