发表机构
Halmstad University; University of Balearic Islands(哈尔姆斯塔德大学; 巴利阿里群岛大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究评估CLIP在零样本全脸和眼周性别估计中的性能,发现全脸准确率高,眼周存在性别偏差,通过阈值对齐和线性SVM可部分改善,但仍逊于全脸表现。
AI 中文摘要
我们研究了CLIP在从全脸和眼周图像进行零样本性别估计中的应用。在Adience数据集的11,299张正面图像上,使用图像-文本相似度与男性/女性提示词评估了三种CLIP骨干网络,无需任务特定训练即可达到95.54%的全脸准确率。对于眼周区域,零样本预测强烈偏向男性,主要原因是决策边界未对齐。阈值对齐显著减少了这种偏差,达到85.29%的准确率。在CLIP特征上训练的线性支持向量机仅带来边际提升,最佳眼周准确率为86.17%,比文献中先前Adience结果高出约2.8%。然而,与全脸性能的差距证实了眼周性别估计的更大难度。
英文摘要
We investigate CLIP for zero-shot gender estimation from full-face and periocular images. Three CLIP backbones are evaluated on 11,299 frontal images from Adience using image-text similarity with male/female prompts, achieving 95.54% full-face accuracy without task-specific training. For periocular, zero-shot predictions are strongly biased towards males, primarily due to a misaligned decision boundary. Threshold alignment substantially reduces this bias, reaching 85.29% accuracy. Linear SVMs trained on CLIP features provide only marginal gains, with a best periocular accuracy of 86.17%, approximately 2.8% above previous Adience results in the literature. Nevertheless, the gap with full-face performance confirms the greater difficulty of periocular gender estimation
CommentsAccepted for publication at 25th International Conference of the Biometrics Special Interest Group, BIOSIG 2026