arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

MS-MFAD:用于人脸活体检测的多模态大语言模型

MS-MFAD : Multimodal large language models for Face Anti-spoofing Detection

Xiaoyong Yu, Rongzhen Li, Shuming Shi, Xinge You

arXiv 2608.17328首次发表:更新:

发表机构

Mashang Consumer Finance Co., Ltd.; Huazhong University of Science and Technology(马上消费金融股份有限公司; 华中科技大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对人脸活体检测的复合威胁与现有方法瓶颈,提出基于多模态大语言模型MFAD的可解释系统,通过细粒度像素-语义锚定机制与少量语义标注实现优异检测性能,鲁棒性与效率更优。

AI 中文摘要

面部生物识别系统当前面临生成式AI与高保真物理欺骗交织的复合威胁,现有防御方法存在泛化能力差、推理不可审计、依赖大量低质量数据集等系统性瓶颈。为解决这些挑战,本文提出用于人脸活体检测的多模态大语言模型MFAD,这是一个面向统一人脸活体检测(UFAD)的可解释推理系统,配套语义级标注基准。与依赖外部工具或粗粒度对齐的方法不同,MFAD通过细粒度像素-语义锚定机制激活多模态大语言模型(MLLMs)的内在推理能力,消除定位幻觉并确保推理路径可审计。本文引入跨攻击的语义级统一标注范式:仅对每种攻击类别标注1000个精确掩码,即可生成与欺骗区域严格对应的推理证据链。在Qwen-VL基础模型上进行监督微调后,该系统使用有限高质量样本,实现域内ACER相对降低40-50%,跨域性能下降控制在11.62%/5.23%以内,显著优于现有框架;在白盒对抗攻击下,检测准确率仅下降3.2%,验证了语义锚定相比基于大量短文本数据训练的模型的鲁棒性。领域从业者对推理路径的证据可靠性评分为4.57/5,推理延迟满足实时部署要求。这些结果证实,少样本高质量语义标注范式可有效构建可信、可解释且成本高效的UFAD系统。

英文摘要

Facial biometric recognition systems currently face compound threats intertwining generative AI and high-fidelity physical spoofing. Existing defenses suffer from systemic bottlenecks, including poor generalization, non-auditable reasoning, and reliance on massive, low-quality datasets. To address these challenges, we propose Multimodal Large Language Models (MFAD) for face anti-spoofing detection, an explainable reasoning system for Unified Face Anti-Spoofing Detection (UFAD), accompanied by a semantic-level annotation benchmark. Unlike methods relying on external tools or coarse alignment, MFAD activates the intrinsic reasoning capabilities of Multimodal Large Language Models (MLLMs) via a fine-grained pixel-semantic anchoring mechanism. This eliminates localization hallucinations and ensures auditable reasoning paths. We introduce a cross-attack semantic-level unified annotation paradigm: by annotating only 1,000 precise masks per attack category, we generate reasoning evidence chains strictly corresponding to spoofed regions. Supervised fine-tuning on the Qwen-VL foundation model demonstrates that, using limited high-quality samples, the system achieves a 40-50% relative reduction in in-domain ACER and restricts cross-domain performance degradation to within 11.62%/5.23%, significantly outperforming existing frameworks. Furthermore, under white-box adversarial attacks, detection accuracy drops by only 3.2%, validating the robustness of semantic anchoring compared to models trained on massive short-text data. Domain practitioners rated the evidence reliability of reasoning paths at 4.57/5, with inference latency satisfying real-time deployment requirements. These results confirm that a few-shot, high-quality semantic annotation paradigm is effective for building trustworthy, explainable, and cost-efficient UFAD systems.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑