arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

CLIP零样本人脸与眼周区域性别估计基准测试

Benchmarking CLIP for Zero-Shot Face and Periocular Gender Estimation

Fernando Alonso-Fernandez, Kevin Hernandez-Diaz, Jose Maria Buades, Josef Bigun

arXiv 2610.06102首次发表:更新:

发表机构

Halmstad University; University of Balearic Islands(哈尔姆斯塔德大学; 巴利阿里群岛大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究评估CLIP在零样本全脸和眼周性别估计中的性能,发现全脸准确率高,眼周存在性别偏差,通过阈值对齐和线性SVM可部分改善,但仍逊于全脸表现。

AI 中文摘要

我们研究了CLIP在从全脸和眼周图像进行零样本性别估计中的应用。在Adience数据集的11,299张正面图像上,使用图像-文本相似度与男性/女性提示词评估了三种CLIP骨干网络,无需任务特定训练即可达到95.54%的全脸准确率。对于眼周区域,零样本预测强烈偏向男性,主要原因是决策边界未对齐。阈值对齐显著减少了这种偏差,达到85.29%的准确率。在CLIP特征上训练的线性支持向量机仅带来边际提升,最佳眼周准确率为86.17%,比文献中先前Adience结果高出约2.8%。然而,与全脸性能的差距证实了眼周性别估计的更大难度。

英文摘要

We investigate CLIP for zero-shot gender estimation from full-face and periocular images. Three CLIP backbones are evaluated on 11,299 frontal images from Adience using image-text similarity with male/female prompts, achieving 95.54% full-face accuracy without task-specific training. For periocular, zero-shot predictions are strongly biased towards males, primarily due to a misaligned decision boundary. Threshold alignment substantially reduces this bias, reaching 85.29% accuracy. Linear SVMs trained on CLIP features provide only marginal gains, with a best periocular accuracy of 86.17%, approximately 2.8% above previous Adience results in the literature. Nevertheless, the gap with full-face performance confirms the greater difficulty of periocular gender estimation

CommentsAccepted for publication at 25th International Conference of the Biometrics Special Interest Group, BIOSIG 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑