发表机构
Idiap Research Institute; University of Lausanne (UNIL)(Idiap研究所; 洛桑大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究在MCIO基准上评估24种冻结基础模型编码器的线性探测性能,发现其仅通过线性分类器可实现强劲的数据集内PAD性能,但跨数据集迁移能力有限,需明确适配以解决域偏移问题。
AI 中文摘要
人脸呈现攻击检测(PAD)在跨数据集评估下仍具挑战性,域偏移会降低在单一数据集上训练的模型性能。大规模标注数据的稀缺促使人们采用预训练视觉模型,而非从头开始训练特定任务架构,这引发了一个根本问题:通用视觉基础模型是否编码了可通过最少特定任务训练获取的PAD相关信息?为探究此问题,我们采用统一的线性探测协议,在MCIO基准(包含MSU-MFSD、CASIA-FASD、Replay-Attack、OULU-NPU)上系统评估24种冻结编码器,涵盖自监督视觉Transformer、视觉-语言编码器及监督CNN。骨干网络保持固定,仅训练轻量级线性头以分离预训练表示中已存在的PAD信息。我们报告了与两个专用PAD基线相比的数据集内和跨数据集性能,以及准确率-计算权衡。结果显示,冻结的基础模型表示仅通过线性分类器即可支持强劲的数据集内PAD性能,但该性能无法可靠地跨数据集迁移。在若干模型族内,模型规模有益,不过该效应非单调,且受架构和预训练的强烈调节。InternViT-6B实现了最低的平均数据集内误差,而CLIP ViT-B/32在评估的探测器中提供了最有利的跨数据集迁移-计算权衡。这些发现表明,尽管预训练表示包含PAD相关信息,但仍需明确的适配以解决域偏移问题。
英文摘要
Face presentation attack detection (PAD) remains challenging under cross-dataset evaluation, where domain shift degrades models trained on a single dataset. The scarcity of large-scale labeled data motivates adapting pretrained vision models rather than training task-specific architectures from scratch, raising a fundamental question: do general-purpose vision foundation models encode PAD-relevant information accessible with minimal task-specific training? To investigate, we systematically evaluate 24 frozen encoders, including self-supervised vision transformers, vision-language encoders, and supervised CNNs, using a unified linear-probing protocol on the MCIO benchmark (MSU-MFSD, CASIA-FASD, Replay-Attack, OULU-NPU). The backbone remains fixed, and only a lightweight linear head is trained to isolate the PAD information already present in the pretrained representation. Results show that frozen foundation-model representations can support strong intra-dataset PAD performance with only a linear classifier, but this performance does not reliably transfer across datasets. Model scale is beneficial within several families, although the effect is not monotonic and is strongly mediated by architecture and pretraining. InternViT-6B achieves the lowest mean intra-dataset error, whereas CLIP ViT-B/32 offers the most favorable cross-dataset transfer-compute trade-off among the evaluated probes. These findings suggest that while pretrained representations contain PAD-relevant information, explicit adaptation remains necessary to address domain shift.
Commentsaccepted at IJCB 2026