发表机构
Michigan State University(密歇根州立大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究通过预计算线性变换将人脸识别模型的人脸嵌入与基础模型对齐,实现了人脸嵌入的自然语言读取、人脸图像渲染及姓名转换,提升了人脸表示的可解释性等能力。
AI 中文摘要
现代人脸识别(FR)的成功很大程度上归功于深度神经网络,这些网络学会从人脸图像中提取紧凑的身份嵌入。这些模型通常针对身份判别进行训练,生成的嵌入对生物特征匹配非常有效,但在语义解释方面大多是不透明的。相比之下,在广泛的视觉或视觉-语言任务上预训练的基础模型,为描述、检索、生成和组织视觉内容提供了丰富的接口。这种对比提出了一个自然的问题:当来自特定领域FR模型的人脸嵌入与基础模型实现互操作时,会获得哪些能力?基于近期关于跨模型嵌入兼容性的研究,我们使用仅从配对嵌入估计的简单预计算线性变换,将现有FR模型与现成的基础模型连接起来。一旦与基础模型对齐,人脸嵌入无需训练或修改任何一个模型,就能以多种方式被“揭示”:它可以用自然语言读取,支持对FR嵌入库进行自由形式的文本查询;可以用未修改的扩散解码器渲染成恢复人物外观的人脸图像;还可以转换为姓名,即使在没有注册人脸库的情况下也能实现识别。实际上,一次线性变换就能将身份嵌入转化为适用于网络规模基础模型的丰富嵌入。这种互操作性将人脸嵌入暴露为具有丰富语义和视觉信息的生物特征表示,对可解释性、检索、重建和模板安全性具有直接影响。
英文摘要
Modern face recognition (FR) owes much of its success to deep neural networks that learn to extract compact identity embeddings from face images. These models are typically trained for identity discrimination, producing embeddings that are highly effective for biometric matching but largely opaque to semantic interpretation. In contrast, foundation models, pretrained on broad visual or vision--language tasks, provide rich interfaces for describing, retrieving, generating, and organizing visual content. This contrast raises a natural question: what capabilities become available when face embeddings from domain-specific FR models are made interoperable with foundation models? Building on recent work on embedding compatibility across models, we use simple pre-computed linear transformations, estimated from paired embeddings alone, to connect existing FR models with off-the-shelf foundation models. Once aligned with a foundation model, a face embedding can be 'unmasked' in multiple ways, without training or modifying either model: it can be read in natural language, enabling free-form text queries over a gallery of FR embeddings; rendered into a face image that recovers a person's appearance, using an unmodified diffusion decoder; and converted to a name, enabling identification even in the absence of an enrolled face gallery. In effect, one linear transformation turns an identity embedding into a rich embedding for web-scale foundation models. This interoperability exposes face embeddings as semantically and visually rich biometric representations, with direct implications for interpretability, retrieval, reconstruction, and template security.