arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

DINO-VPT:用于联合物理-数字人脸反欺骗的分层视觉提示调整

DINO-VPT: Hierarchical Visual Prompt Tuning for Joint Physical-Digital Face Anti-Spoofing

Pierre Gallin-Martel, Mika Feng, Koichi Ito, Takafumi Aoki

arXiv 2607.20900首次发表:更新:

发表机构

Graduate School of Information Sciences, Tohoku University(东北大学信息科学研究生院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究针对人脸反欺骗需求,提出DINO-VPT轻量级仅视觉框架,利用分层视觉提示调整,通过提示路由网络动态注入提示,无需多模态融合,在基准测试中比现有方法准确率更高,证明合理架构能达最优性能。

AI 中文摘要

随着欺骗攻击的日益多样化,对能够检测物理和数字威胁的统一人脸反欺骗(FAS)模型的需求不断增长。虽然现有的视觉语言模型(VLM)在此背景下具有较高的泛化能力,但它们严重依赖复杂的多模态融合和外部文本编码器。本文提出了DINO-VPT,一个轻量级的、仅视觉的框架,利用分层视觉提示调整。通过提示路由网络(PRN)根据输入特征动态注入提示,该方法无需多模态融合就能有效解开各种欺骗伪像。在UniAttackData基准上的评估表明,DINO-VPT比基于VLM的现有方法具有更高的准确率。结果表明,结构合理的仅视觉架构无需多模态监督就能在统一FAS中达到最优性能。

英文摘要

With the increasing diversity of spoofing attacks, there is a growing demand for unified Face Anti-Spoofing (FAS) models capable of detecting both physical and digital threats. While existing Vision-Language Models (VLMs) demonstrate high generalization in this context, they heavily rely on complex multimodal fusion and external text encoders. In this paper, we propose DINO-VPT, a lightweight, vision-only framework leveraging hierarchical visual prompt tuning. By dynamically injecting prompts conditioned on input features via a Prompt Routing Network (PRN), our method effectively disentangles diverse spoofing artifacts without requiring multimodal fusion. Evaluations on the UniAttackData benchmark demonstrate that DINO-VPT achieves higher accuracy than state-of-the-art VLM-based methods. Our results indicate that a properly structured vision-only architecture can achieve state-of-the-art performance in unified FAS without the need for multimodal supervision.

Commentsaccepted to IJCB2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑