arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.21486cs.CV

EXPL-FR:基于视觉-语言对齐的人脸识别模型解释方法

EXPL-FR: Explaining Face Recognition Models via Vision-Language Alignment

Guray Ozgur, Mustafa Efe Tamyapar, Naser Damer, Fadi Boutros

首次发表
浏览论文内容

中文总结 AI 辅助

EXPL-FR是一种基于视觉-语言对齐的人脸识别模型解释方法,通过轻量适配器在FR嵌入空间内生成语义特征,支持多粒度解释,可实现无标签的属性级审计与模型性能排名。

中文摘要 AI 辅助

深度人脸识别(FR)模型的准确率已接近饱和,但仍存在可解释性不足的问题:从业者无法获知相似度得分依赖于哪些语义属性。EXPL-FR在FR模型自身的嵌入空间内解决了这一问题。一种轻量适配器将视觉-语言模型(VLM)的图像编码器与冻结的FR空间对齐,该适配器仅在人脸图像上训练,从未使用文本数据。由于VLM的编码器共享同一空间,相同的适配器也适用于文本编码器,可将22个类别的978个属性提示(还可扩展)转化为FR空间锚点,且无需额外成本。我们并未假设这种迁移必然有效:通过人脸验证协议对其进行评估,并通过仅改变适配器的消融实验来分离其贡献。并非所有概念都能迁移,因为FR模型为了在跨身份验证时保持不变性,会丢弃一些因素。一种无标签可检测性指标将每个概念在FR空间中的可分离性与VLM空间中的可分离性进行比较,其中100个最具可检测性的概念构成了模型的可读语义特征,该特征在身份区分方面优于完整词汇表。我们覆盖了4种FR骨干网络和2种VLM编码器,EXPL-FR无需访问模型架构,支持身份级、单图像级和差异级解释。我们在三种监督设置(人类标签(当前实践)、VLM伪标签、我们的全提示驱动审计)下对属性级审计进行了基准测试,与真实验证行为进行对比。在无标签的情况下,提示驱动审计可按测量的FR模型的每种族RFW误差对4个FR模型进行排名,并按真实验证成本对受控属性变化进行排名。

英文摘要

Deep face recognition (FR) models reach near-saturated accuracy but remain opaque: a practitioner cannot ask which semantic attributes a similarity score relied upon. EXPL-FR answers this inside the FR model's own embedding space. A lightweight adapter aligns a vision-language model's (VLM) image encoder with the frozen FR space, trained on face images alone and never on text. Because the VLM's encoders share one space, the same adapter applies to the text encoder, turning 978 attribute prompts in 22 categories, also extendable, into FR-space anchors at no extra cost. We do not assume this transfer works: a face-verification protocol measures it, and an ablation changing only the adapter isolates its contribution. Not every concept survives, because an FR model earns its invariances by discarding the factors it must verify identities across. A label-free detectability measure compares each concept's separability in FR space against the VLM space, and the 100 most detectable form the model's readable semantic signature, which separates identities better than the full vocabulary. We cover four FR backbones and two VLM encoders, EXPL-FR needs no architecture access, and supports identity-level, per-image, and differential explanations. We benchmark attribute-level auditing under three supervision settings, human labels (current practice), VLM pseudo-labels, and our fully prompt-driven audit, against real verification behavior. With no labels, the prompt-driven audit ranks four FR models by their measured per-ethnicity RFW errors and ranks controlled attribute changes by their true verification cost.

发表机构

  • Fraunhofer IGD(弗劳恩霍夫应用研究促进协会图形数据处理研究所)
  • TU Darmstadt(达姆施塔特工业大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑