arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

用于可解释人脸识别的视觉语言模型融合

Vision Language Model Fusion for Explainable Face Recognition

Ana Estrada-Real, Lydia Alapatt, Christoph Busch, Christian Rathgeb

arXiv 2608.24430首次发表:更新:

发表机构

Hochschule Darmstadt(达姆施塔特应用科学大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究提出多视觉语言模型(VLM)融合框架,结合相似度分数、文本解释与人脸图像,提升了人脸识别准确率,同时生成更丰富稳健的可解释决策,助力开发负责任的可解释人脸识别系统。

AI 中文摘要

人脸识别系统的负责任部署不仅需要准确的决策,还需提供可解释、可审计的证据,使用户能够理解、评估并质疑其决策。视觉语言模型(VLM)结合视觉分析与自然语言推理,为可解释人脸识别提供了有前景的基础,但依赖单一模型可能会限制决策准确性及所提供的解释。本研究探讨是否可通过融合多个VLM来提升识别准确率并丰富决策相关解释。首先评估4种VLM作为独立人脸识别系统的性能,随后提出一种融合框架:两个源模型提供相似度分数与文本解释,第三个VLM作为决策模型。设置4种不同融合场景,逐步向决策模型提供分数、解释、人脸图像及这些模态的组合。整体研究发现,多VLM融合的价值不仅限于识别性能:VLM可提供互补的解释与视角,使人脸识别决策的解释更丰富,提升透明度、可审计性与错误分析能力,这对开发负责任的可解释人脸识别系统至关重要,用户与操作者不仅需理解最终决策,还需了解其背后的证据与潜在来源。所提出的融合决策分数、解释与人脸图像的多模态VLM,识别准确率优于现有最优VLM及领域特定人脸识别模型,且融合解释预计比单个VLM生成的解释更稳健。

英文摘要

Responsible deployment of face verification systems requires more than accurate decisions: systems should also provide interpretable and auditable evidence that enables users to understand, assess, and challenge their decisions. Vision-language models (VLMs) provide a promising foundation for explainable face recognition by combining visual analysis with natural-language reasoning. However, relying on a single model may further limit the decision accuracy as well as provided explanations. This work therefore investigates whether multiple VLMs can be combined to improve recognition accuracy, and to enrich the explanations associated with those decisions. This work evaluates four VLMs as standalone face verification systems and subsequently proposes a fusion framework, where two source models provide similarity scores and textual justifications and a third VLM acts as a decider model. Four different fusion scenarios are considered, progressively providing the decider model with scores, justifications, face images, and combinations of these modalities. Overall, the findings suggest that the value of multi-VLM fusion extends beyond recognition performance. VLMs can provide complementary justifications and perspectives that enable richer explanations of face recognition decisions, supporting greater transparency, auditability, and error analysis. This is relevant to the development of responsible explainable face verification systems, where users and operators should be able to understand not only the final decision but also the evidence and potential sources underlying it. The proposed multimodal VLM, which combines decision scores, explanations, and face images, achieves higher recognition accuracy than state-of-the-art VLMs and domain-specific face recognition models, while also providing fused explanations that are expected to be more robust than those generated by individual VLMs.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑