arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AlignFace:具有可解释概念关系的人类对齐人脸相似度度量

AlignFace: Human-Aligned Face Similarity Metric with Interpretable Concept Relations

Ying Huang, Wencan Zhang, Brian Y. Lim

arXiv 2608.14130首次发表:更新:

发表机构

National University of Singapore(新加坡国立大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究针对人脸生成模型评估中现有度量未实现认知对齐的问题,提出AlignFace度量,结合VLM、CA、CBM、GAM等技术,利用FACETS数据集,提升了与人类亚群感知的对齐度。

AI 中文摘要

用于生成人脸内容的计算机视觉模型,如人脸编辑和隐私保护,正日益影响人们,需要能忠实反映人类感知的相似度度量。尽管感知评估已从基于信号的启发式方法发展到基于表征的度量,但现有方法局限于行为建模,未实现认知对齐。它们依赖隐含且虚假的关系,假设存在通用观察者,未考虑不同人群的固有差异,导致利益相关者的评估模型不准确,给生成模型调试带来误导。我们未将感知视为黑箱,而是利用人类人脸相似度感知的认知心理学研究成果:依赖人脸特征和构型属性、非线性心理物理响应缩放、内群体偏差。我们推出FACETS数据集,并提出AlignFace——一种可解释、人类对齐的人脸相似度度量,通过事前建模编码这些认知原则。它采用视觉语言建模(VLM)编码成对人脸图像和基于文本的属性,门控交叉注意力(CA)提取属性特定的人脸差异表征,概念瓶颈建模(CBM)通过可解释人脸属性约束推理,神经广义加性模型(GAM)建模其非线性影响。实验发现,与基线度量(包括近期无领域学习的感知度量)相比,AlignFace显著提升了与人类亚群感知的对齐度。通过连接学习表征与人类认知过程,本研究为人脸图像提供了更透明、对齐的感知评估度量。

英文摘要

Computer vision models for generated facial content, such as face editing and privacy protection, increasingly affect people, requiring similarity metrics that serve as faithful proxies for human perception. While perceptual evaluation has progressed from signal-based heuristics to representation-based metrics, current approaches are limited to behavioral modeling without cognitive alignment. They rely on implicit and spurious relations while assuming a universal observer, failing to account for inherent variations across diverse human populations. This leads to inaccurate evaluative models of stakeholders and misleading guidance for generative model debugging. Rather than treating perception as a black box, we leverage scientific findings from cognitive psychology of human face similarity perception: dependence on facial featural and configural attributes, nonlinear psychophysical response scaling, and own-group biases. We introduce the FACETS dataset and propose AlignFace, an interpretable, human-aligned, face similarity metric that encodes these cognitive principles through ante-hoc modeling. It employs visual-language modeling (VLM) to encode paired face images and text-based attributes, gated cross-attention (CA) to extract attribute-specific facial difference representations, concept bottleneck modeling (CBM) to constrain reasoning via interpretable face attributes, and neural generalized additive model (GAM) to model their nonlinear influence. Experiments found AlignFace significantly improves alignment with human subpopulation perceptions compared to baseline metrics, including recent domain-free learned perceptual metrics. By bridging learned representations and human cognitive processes, this work enables more transparent and aligned perceptual evaluation metrics for face images.

Comments10 pages, 10 figures, 2 tables, ACM MM 26

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑