arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

FAS-R1:用于推理人脸防伪的统一多任务多模态大语言模型(MLLM)

FAS-R1: A Unified Multi-Task MLLM for Reasoning Face Anti-Spoofing

Hongyang Wang, Yichen Shi, Hongrui Li, Yiru Huo, Jun Feng, Zitong Yu

arXiv 2607.26432首次发表:更新:

AI 中文总结

本文提出面向推理的两阶段多任务MLLM框架FAS-R1,结合DSA与DA-GRPO优化,在人脸防伪的分类、识别、定位任务中表现优异,性能优于对比系统且具备良好缩放性。

AI 中文摘要

人脸防伪(FAS)不仅需要做出真实/伪造判断,还需提供攻击语义和图像层面证据供人工检查。现有判别式FAS模型多以标签为中心,而近期基于多模态大语言模型(MLLM)的方法虽能生成结构化输出,但仍主要依赖监督微调,常产生模板化理由且对难攻击的优化不足。本文提出FAS-R1,这是一种面向推理的两阶段统一FAS预测MLLM框架,涵盖真实性分类、攻击类型识别和伪造区域定位。FAS-R1首先使用高质量长思维链数据集FAS-R1-23K进行冷启动监督微调,随后执行FAS专用的广义偏好优化(GRPO)后训练。退化模拟增强(DSA)可鼓励模型在视觉质量变化下稳定推理伪造线索,而难度感知GRPO(DA-GRPO)可缓解易样本主导问题,避免难任务-攻击组(尤其是化妆、面具等微妙或模糊攻击)优化不足。3B规模的FAS-R1主模型在域内实现98.75%的真实性准确率、93.33%的攻击类型准确率,以及96.30/94.73%的AP@40/AP@50指标,在跨域真实性泛化和答案与理由质量上均优于对比系统,不同基础模型的实验也显示出良好的缩放特性,代码即将发布。

英文摘要

Face anti-spoofing (FAS) is increasingly expected to provide not only bona fide/spoof decisions, but also attack semantics and image-grounded evidence for human inspection. Existing discriminative FAS models remain largely label-centric, while recent MLLM-based methods offer structured outputs but still rely mainly on supervised fine-tuning, often producing template-like rationales and weak optimization for difficult attacks. We propose FAS-R1, a two-stage reasoning-oriented MLLM framework for unified FAS prediction, covering authenticity classification, attack-type recognition and spoof-region localization. FAS-R1 first uses FAS-R1-23K, a high-quality long-CoT dataset, for cold-start supervised fine-tuning, and then performs FAS-specific GRPO post-training. Degradation-Simulated Augmentation (DSA) encourages stable spoof-cue reasoning across visual-quality shifts, while Difficulty-Aware GRPO (DA-GRPO) mitigates easy-sample dominance that may leave difficult task--attack groups under-optimized, especially for subtle or ambiguous attacks such as makeup and mask attacks. The main 3B FAS-R1 model achieves 98.75\% authenticity accuracy, 93.33\% attack-type accuracy, and 96.30/94.73\% AP@40/AP@50 in-domain. It also outperforms the compared systems in cross-domain authenticity generalization and answer-and-rationale quality. Experiments with different base models further show favorable scaling behavior. The code will be released soon.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑