arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.17351cs.CV

面向开放世界人脸活体检测的基元驱动组合式取证视觉提示

Primitive-Driven Compositional Forensic Visual Prompting for Open-World Face Anti-Spoofing

Fangling Jiang, Qi Li, Bing Liu, Weining Wang, Quilin Huang, Zhenan Sun, Ming-Hsuan Yang

首次发表
浏览论文内容

中文总结 AI 辅助

针对开放世界人脸活体检测的协变量与语义偏移问题,提出基元驱动的组合式取证视觉提示框架,在9个协议上实现先进性能,具备强跨域泛化与未见攻击适配能力。

中文摘要 AI 辅助

开放世界人脸活体检测必须同时应对协变量偏移和语义偏移:源域与目标域在成像条件上存在差异,而目标域包含训练中未出现的各类攻击类型。现有基于提示的方法常通过类别语义或语言指导来表达伪造,这对建模高层概念有效,但难以显式捕捉未见攻击不断演变的细粒度且空间异质的取证证据。受“许多未见攻击可通过重复视觉线索的新组合来表征”这一假设的启发,我们提出一种完全在视觉特征空间中运行的组合式取证视觉提示学习框架。在基于冻结ViT的视觉基础模型上,该框架采用 patch-aware 注意力将一组可学习的微取证基元优化为源自图像块的局域取证证据单元;特定类别的全局上下文提示则提供依赖输入的路由权重,自适应选择并组合这些基元,形成用于真实/伪造判别的组合式取证视觉提示。这些基元未被赋予预定义语义,其专门化与复用源于跨数据集的共享参数化与联合优化。在9个开放世界协议上的实验表明,该方法达到了先进性能,具备强跨域泛化能力,且对未见攻击的适配性良好。

英文摘要

Open-world face anti-spoofing must address both covariate and semantic shifts: source and target domains differ in imaging conditions, while target domains contain diverse attack types absent from training. Existing prompt-based approaches often express spoofing through category semantics or language guidance, which is effective for modeling high-level concepts but is less suited to explicitly capturing the evolving fine-grained and spatially heterogeneous forensic evidence of unseen attacks. Motivated by the hypothesis that many unseen attacks can be characterized by new combinations of recurring visual cues, we propose a compositional forensic visual prompt learning framework that operates entirely in the visual feature space. Built on a frozen ViT-based vision foundation model, the framework employs patch-aware attention to refine a shared set of learnable micro-forensic primitives into localized forensic evidence units derived from image patches. Class-specific global contextual prompts then provide input-dependent routing weights that adaptively select and compose these primitives into compositional forensic visual prompts for real/spoof discrimination. The primitives are not assigned predefined semantic meanings; instead, their specialization and reuse emerge from shared parameterization and joint optimization across categories. Extensive experiments on nine open-world protocols demonstrate state-of-the-art performance, strong cross-domain generalization, and robust adaptation to unseen attacks.

发表机构

  • School of Computer Science, University of South China(南华大学计算机学院)
  • MAIS, CASIA(中国科学院自动化研究所模式识别国家重点实验室)
  • University of California, Merced(加州大学默塞德分校)

机构由 AI 辅助整理,请以论文原文为准。

↑