arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

用于人脸呈现与变形攻击检测的基础模型及多模态大语言模型

Foundation and Multimodal Large Language Models for Face Presentation and Morph Attack Detection

Hatef Otroshi Shahreza, Asif Hussain Khan, Peter Lorenz, Alain Komaty, Sébastien Marcel

arXiv 2608.29802首次发表:更新:

发表机构

Idiap Research Institute; Université de Lausanne(Idiap研究所; 洛桑大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文研究通用基础模型和多模态大语言模型是否包含人脸呈现与变形攻击相关信息,通过五种方法在8个公开数据集上测试,发现微调后模型跨数据集检测性能达SOTA,将开源代码。

AI 中文摘要

人脸识别系统正越来越多地部署到对安全性要求极高的应用场景中,但它们仍然容易受到呈现攻击和变形攻击的影响。因此,呈现攻击检测(PAD)和变形攻击检测(MAD)是可信人脸生物识别技术的必要组成部分。尽管PAD和MAD方法取得了一定进展,但现有检测器的泛化能力有限,在跨数据集评估中性能会下降。本文系统研究了通用基础模型(FMs)和多模态大语言模型(MLLMs)是否编码了与PAD和MAD相关的信息,以及如何最好地将这些模型部署到这两项任务中。我们研究了五种对模型内部信息访问程度递增的方法:(i)对现成MLLMs的零样本提示;(ii)在MLLM输出的下一个token对数概率上训练一个浅层模型;(iii)对任务特定的问答数据进行参数高效微调,得到两个专门的MLLM,分别称为PADLLM和MADLLM,它们还为自身的决策提供文本推理;(iv)对冻结的视觉编码器进行线性探测;(v)对FMs和MLLMs的视觉编码器进行微调。我们在四个PAD数据集(MSU-MFSD、CASIA-FASD、Replay-Attack和OULU-NPU)和四个MAD数据集(FFHQ、FRGC、FRLL和FERET)上对16个开放权重MLLMs和30个视觉编码器骨干进行了基准测试。实验表明,FMs和MLLMs在PAD和MAD任务中可实现显著性能;此外,微调后的模型在跨数据集评估中达到了最先进的检测性能,表明通用预训练表示包含大量与攻击相关的信息。我们所有实验的源代码将公开发布。

英文摘要

Face recognition systems are increasingly deployed in security-critical applications, yet they remain vulnerable to presentation and morph attacks. Presentation attack detection (PAD) and morphing attack detection (MAD) are therefore essential components of trustworthy face biometrics. Despite advancements in PAD and MAD methods, existing detectors suffer from limited generalization and degrade in cross-dataset evaluation. In this paper, we systematically investigate whether general-purpose foundation models (FMs) and multimodal large language models (MLLMs) encode PAD-relevant and MAD-relevant information, and how such models can best be deployed for both tasks. We study five approaches with increasing access to the internal information of the model: (i) zero-shot prompting of off-the-shelf MLLMs; (ii) training a shallow model on the next-token logit probabilities at the output of the MLLM; (iii) parameter-efficient fine-tuning on task-specific question-answer data, yielding two specialized MLLMs, called PADLLM and MADLLM, which additionally provide textual reasoning for their decisions; (iv) linear probing of frozen vision encoders; and (v) fine-tuning of vision encoders of FMs and MLLMs. We benchmark 16 open-weight MLLMs and 30 vision encoder backbones on four PAD datasets (MSU-MFSD, CASIA-FASD, Replay-Attack, and OULU-NPU) and four MAD datasets (FFHQ, FRGC, FRLL, and FERET). Our experiments show that FMs and MLLMs can achieve significant performance for PAD and MAD. In addition, the fine-tuned models achieve state-of-the-art detection performance in cross-dataset evaluation, indicating that general-purpose pretrained representations carry substantial attack-relevant information. Source code of all our experiments will be publicly released.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑