发表机构
Idiap Research Institute; Université de Lausanne(伊迪亚普研究所; 洛桑大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出一个基准测试框架,将解释质量作为核心评估轴,通过相关性和忠实性两个标准量化开放权重视觉语言模型在人脸识别中的解释,并联合评估验证准确性与解释质量,揭示现有不足。
AI 中文摘要
视觉语言模型(VLMs)最近被提出作为人脸识别的有前景工具,因为它们能够生成自然语言解释以及相似度分数。这一能力被认为在法医背景下的人脸比较中具有吸引力,因为此类比较要求决策透明且可审计。然而,现有的针对该用例的VLM评估主要集中在识别准确性上,而生成解释的有效性仍未量化。在这项工作中,我们引入了一个基于VLM的人脸识别基准测试框架,将解释质量作为核心评估轴。我们提出了解释应满足的两个标准:相关性,即依赖于身份稳定的面部特征;以及忠实性,即与可见图像内容一致且不包含幻觉特征。我们联合开发了一种方法论,通过将模型输出限制为支持自动化查询和审计的结构化解释格式,来量化所评估模型的相关性和忠实性。利用该框架,我们对多个开放权重VLM系列进行了基准测试,联合评估了人脸验证准确性和解释质量。我们的结果突显了生成解释仍存在的不足,并强调了需要此类解释质量指标以全面了解模型性能。所提出的基准测试和开源评估工具为可解释人脸识别系统的正确基准测试和未来微调提供了基础。
英文摘要
Vision-Language Models (VLMs) have recently been proposed as promising tools for face recognition, as they can produce natural language explanations alongside similarity scores. This capability is considered appealing for face comparisons in forensic contexts, which require decisions to be transparent and auditable. However, existing evaluations of VLMs for that use case focus mostly on recognition accuracy, while the validity of generated explanations remains unquantified. In this work, we introduce a benchmarking framework for VLM-based face recognition that treats explanation quality as a core evaluation axis. We propose two criteria that explanations should satisfy: relevance, i.e., reliance on identity-stable facial features; and faithfulness, i.e., alignment with the visible image content without hallucinated features. We jointly develop a methodology enabling the quantification of relevance and faithfulness of evaluated models, based on constraining model outputs to a structured explanation format that supports automated querying and auditing. Using this framework, we benchmark several families of open-weight VLMs, jointly evaluating face verification accuracy and explanation quality. Our results highlight remaining shortcomings of produced explanations, and emphasize the need for such explanation quality metrics to get a complete picture of model performance. The proposed benchmark and open-source evaluation harness provide a foundation for proper benchmarking and future fine-tuning of explainable face recognition systems.
Comments11 pages