发表机构
University of Wisconsin–Green Bay(威斯康星大学格林湾分校)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究四指SLAP指纹验证,介绍SLAPBench基准,评估多个MLLM在不同提示下的表现,发现提示控制崩溃,模型能力控制歧视,建立了特定于SLAP的MLLM基线,揭示了模型在指纹验证中的能力差距和公平性问题。
AI 中文摘要
四指SLAP指纹是单手食指、中指、无名指和小指的平面实时扫描印记,用于边境管制和执法中的身份验证。此前没有基准测试评估多模态大语言模型(MLLM)能否从SLAP图像验证身份。本文介绍了SLAPBench,这是首个基于MLLM的四指SLAP指纹验证基准,由NIST SD302b构建,包含7832对数据(176对匹配,7656对不匹配)。评估了四个开源MLLM和专有模型Claude Opus 4.8在零样本、任务描述和相似度评分提示下的情况。结果表明,任务描述提示会使开源模型的错误接受率接近100%,Gemma - 3 - 12B在零样本时也表现不佳;Claude Opus 4.8在两种二元提示下表现最佳(错误接受率 = 20.2%)。相似度评分揭示了开源模型间的能力差距,Qwen3 - VL - 8B达到完美分离(AUC = 1.000)。公平性探测表明,随着歧视减弱,性别、种族和年龄方面的差异会增大。SLAPBench建立了首个特定于SLAP的MLLM基线,表明提示控制崩溃,而模型能力控制歧视。
英文摘要
Four-finger SLAP fingerprints are flat live-scan impressions of the index, middle, ring, and little fingers of one hand, used for identity verification in border control and law enforcement. No benchmark has evaluated whether multimodal large language models (MLLMs) can verify identity from SLAP images. We introduce SLAPBench, the first benchmark for MLLM-based four-finger SLAP fingerprint verification, built from NIST SD302b with 7,832 pairs (176 mated, 7,656 non-mated). We evaluate four open-source MLLMs (InternVL3-8B, Qwen2.5-VL-7B, Qwen3-VL-8B, Gemma-3-12B) and the proprietary Claude Opus 4.8 under zero-shot, task-description, and similarity-scoring prompts. Prompting governs verification behavior. Task-description prompting collapses all four open-source models to near-100% False Accept Rate (FAR), and Gemma-3-12B collapses under zero-shot as well; Claude Opus 4.8 alone resists collapse under both binary prompts, giving the best binary result (FAR = 20.2%). Similarity scoring removes collapse across the open-source models and exposes wide capability gaps: Claude reaches AUC = 0.953 and Gemma-3-12B 0.837, while InternVL3-8B is inverted (AUC = 0.590) and Qwen2.5-VL-7B near random (0.567). Qwen3-VL-8B attains perfect separation (AUC = 1.000), which we treat as a diagnostic rather than as capability: SD302b holds one SLAP capture per finger position, so mated pairs are cross-resolution. A matched-resolution control leaves the perfect score intact, ruling out the resolution shortcut; what cannot be excluded within SD302b is near-duplicate detection, since a mated pair is one capture rendered twice. A fairness probe over gender, race, and age suggests disparity grows as discrimination weakens. SLAPBench establishes the first SLAP-specific MLLM baseline and shows that prompting governs collapse while model capability governs discrimination.
Comments19 pages, 6 figures, 2 tables. Includes appendix with supporting figures and per-subgroup fairness detail. Code and data: https://github.com/bibeshpyakurel/SLAPBench