发表机构
Aalborg University; Pioneer Centre for AI; Ducaltus Ltd.; Milestone Systems(奥尔堡大学; 先锋人工智能中心; 杜卡尔图斯有限公司; 里程碑系统公司)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出CIFA框架,用于识别人脸分析中人口统计与上下文属性交互导致的隐藏子群漏洞,经多数据集多模型评估发现现有公平性评估会掩盖交叉差异,且单一缓解策略无法持续消除此类差异。
AI 中文摘要
计算机视觉中的公平性评估通常依赖于聚合准确率和人口统计子群分析。然而,视觉模型也对光照、模糊、图像质量、面部配饰及外观属性等上下文因素敏感,这些因素可能与人口统计特征相互作用,形成隐藏子群——即便聚合准确率高、人口统计公平性看似可接受,这些子群的性能仍会大幅下降。为解决该问题,本文提出上下文交叉公平性审计框架(Contextual-Intersectional Fairness Auditing Framework, CIFA),这是一种用于识别人口统计属性与上下文属性相互作用导致的子群漏洞的结构化框架。CIFA依次执行人口统计审计、上下文审计及上下文交叉审计,随后通过最差组发现来识别并排序最脆弱的属性组合。本文在FairFace、CelebA和UTKFace数据集上,使用ResNet-50和ViT-B/16模型对性别分类任务评估CIFA。结果表明,聚合准确率和仅人口统计评估会掩盖显著的上下文交叉差异。本文进一步通过“审计-缓解-重审计”协议评估多种已有的缓解策略,发现尽管部分最差组差异有所降低,但没有任何单一策略能在所有数据集和模型架构上持续消除这些差异。这些发现确立了上下文交叉审计作为公平性评估重要组成部分的地位,并提供了一个可复现的框架,用于发现、优先级排序和重新评估人脸分析系统中的隐藏子群风险。
英文摘要
Fairness evaluation in computer vision commonly relies on aggregate accuracy and demographic subgroup analysis. However, visual models are also sensitive to contextual factors such as illumination, blur, image quality, facial accessories, and appearance attributes. These factors may interact with demographic characteristics, producing hidden subgroups in which performance degrades substantially despite strong aggregate accuracy and apparently acceptable demographic fairness. To address this, we propose the Contextual-Intersectional Fairness Auditing Framework (CIFA), a structured framework for identifying subgroup vulnerabilities arising from interactions between demographic and contextual attributes. CIFA performs demographic, contextual, and contextual-intersectional auditing, followed by worst-group discovery to identify and rank the most vulnerable attribute combinations. We evaluate CIFA on gender classification using ResNet-50 \cite{he2016deep} and ViT-B/16 \cite{dosovitskiy2020image} across FairFace \cite{Karkkainen2021}, CelebA \cite{Liu2015}, and UTKFace \cite{Zhang2017}. Our results show that aggregate accuracy and demographic-only evaluation can mask substantial contextual-intersectional disparities. We further assess several established mitigation strategies through an audit--mitigate--reaudit protocol and find that, although some worst-group disparities are reduced, no single strategy consistently eliminates them across datasets and architectures. These findings establish contextual-intersectional auditing as an important component of fairness evaluation and provide a reproducible framework for discovering, prioritizing, and reassessing hidden subgroup risks in face analysis systems.
Comments16 pages, 2 figures