发表机构
Idiap Research Institute; University of Lausanne (UNIL)(伊迪亚普研究所; 洛桑大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究系统评估32种基础模型,发现LoRA适配仅优化数据集内人脸PAD性能,跨数据集泛化仍受限,指出仅靠LoRA适配不足以解决人脸PAD跨数据集性能下降问题。
AI 中文摘要
人脸呈现攻击检测(PAD)旨在可靠检测各类呈现攻击。尽管PAD方法在单个数据集内表现出色,但在跨数据集评估中性能会下降,传感器或光照条件的差异会使检测器的有效性从近乎完美降至接近随机。基础模型(FMs)成为颇具前景的替代方案,因为典型的PAD数据集(如MCIO基准,包含MSU-MFSD、CASIA-FASD、Replay-Attack和OULU-NPU)相较于用于网络预训练的规模较小。然而,现有PAD系统主要聚焦于基于CLIP的基础模型,却忽略了其他具有不同架构和训练流程的FMs。本研究通过系统评估32种FMs来解决这一问题:零样本提示在所有模型家族和规模下的性能接近随机;当视觉编码器采用低秩适配(LoRA)且可训练权重占比低于1%时,多数情况下数据集内的ACER低于2%,但跨数据集的ACER显著更高;LoRA主要优化数据集内的决策边界,表明预训练表示和适配数据集对跨数据集泛化的作用,大于所评估的轻量适配策略的作用。
英文摘要
Face presentation attack detection (PAD) aims to reliably detect a wide range of presentation attacks. While PAD methods achieve strong performance within individual datasets, their performance degrades under cross-dataset evaluation. Variations in sensors or lighting conditions can reduce the effectiveness of detectors from near-perfect to nearly random. Foundation models (FMs) have emerged as a promising alternative because typical PAD datasets, such as the MCIO benchmarks (MSU-MFSD, CASIA-FASD, Replay-Attack, and OULU-NPU), are small relative to the scale used for web-based pretraining. However, existing PAD systems primarily focus on CLIP-based foundation models, while overlooking other FMs with different architectures and training procedures. This study addresses this question by systematically evaluating 32 FMs. Zero-shot prompting achieves performance near chance across model families and scales. The vision encoders, when low-rankadapted (LoRA) with fewer than 1% trainable weights, achieve below 2% intra-dataset ACER in most cases, while cross-dataset ACER is substantially higher. LoRA primarily refines the decision boundary within a dataset, suggesting that pretrained representations and the adaptation dataset play a larger role in cross-dataset generalization than the evaluated lightweight adaptation strategy.