发表机构
Graduate School of Information Sciences, Tohoku University(东北大学信息科学研究生院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究针对人脸反欺骗需求,提出DINO-VPT轻量级仅视觉框架,利用分层视觉提示调整,通过提示路由网络动态注入提示,无需多模态融合,在基准测试中比现有方法准确率更高,证明合理架构能达最优性能。
AI 中文摘要
随着欺骗攻击的日益多样化,对能够检测物理和数字威胁的统一人脸反欺骗(FAS)模型的需求不断增长。虽然现有的视觉语言模型(VLM)在此背景下具有较高的泛化能力,但它们严重依赖复杂的多模态融合和外部文本编码器。本文提出了DINO-VPT,一个轻量级的、仅视觉的框架,利用分层视觉提示调整。通过提示路由网络(PRN)根据输入特征动态注入提示,该方法无需多模态融合就能有效解开各种欺骗伪像。在UniAttackData基准上的评估表明,DINO-VPT比基于VLM的现有方法具有更高的准确率。结果表明,结构合理的仅视觉架构无需多模态监督就能在统一FAS中达到最优性能。
英文摘要
With the increasing diversity of spoofing attacks, there is a growing demand for unified Face Anti-Spoofing (FAS) models capable of detecting both physical and digital threats. While existing Vision-Language Models (VLMs) demonstrate high generalization in this context, they heavily rely on complex multimodal fusion and external text encoders. In this paper, we propose DINO-VPT, a lightweight, vision-only framework leveraging hierarchical visual prompt tuning. By dynamically injecting prompts conditioned on input features via a Prompt Routing Network (PRN), our method effectively disentangles diverse spoofing artifacts without requiring multimodal fusion. Evaluations on the UniAttackData benchmark demonstrate that DINO-VPT achieves higher accuracy than state-of-the-art VLM-based methods. Our results indicate that a properly structured vision-only architecture can achieve state-of-the-art performance in unified FAS without the need for multimodal supervision.
Commentsaccepted to IJCB2026