arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2610.08279cs.CV

这到底是谁的脸?针对儿童面部情感识别的多模型审计,以及为什么差距在于头部而非特征

Whose Face Is It Anyway? A Multi-Model Audit of Facial Affect Recognition on Children, and Why the Gap Is the Head, Not the Features

  • University of Augsburg(奥格斯堡大学)
  • University of the Bundeswehr Munich(慕尼黑联邦国防军大学)

机构由 AI 辅助整理,请以论文原文为准。

Tobias Hallmen, Robin-Nico Kampa, Elisabeth André

AI总结:

针对儿童面部情感识别,审计五个预训练模型发现性能差距源于分类器头部而非特征,通过仅重校准头部即可显著提升儿童识别性能。

AI中文摘要:

面部情感模型几乎完全在成年人数据上训练,却越来越多地被应用于教育、健康和发展研究中的儿童。我们通过一个共享测试框架,对五个在AffectNet上预训练的表情模型(EmoNet、EmotiEffLib、DDAMFN++、OpenFace 3.0、LibreFace)在儿童数据上进行了受控的多模型审计,覆盖四个儿童图像数据集、AffectNet-8验证集和两个自发性儿童视频数据集。研究得出三个发现。第一,儿童差距与模型无关:每个架构从摆拍面部到自然面部性能均下降,并共享恐惧到惊讶的混淆。第二,该差距在所有五个模型中集中且相互印证:张嘴面部(被识别为惊讶,与AU26下颌张开相关)和南亚儿童系统性性能下降,而回避目光的惩罚较小;闭嘴面部、白人和黑人儿童以及直视目光则没有此问题;偏差与表情形态和特定人群相关,而非肤色。第三,该差距是可诊断的:在冻结特征上的线性探针在未见儿童上达到0.75-0.91,而零样本仅为0.48-0.66,因此差距主要存在于分类器头部而非表示中,而维度效价/唤醒回归在域偏移下急剧下降。基于此,仅用少量目标数据重新校准头部,在两个最大的儿童数据集上所有五个模型恢复+0.13至+0.28的性能,且对成年人影响可忽略,但增益仅在分布内,不能跨儿童数据集迁移。我们将发布测试框架、逐样本预测和分析代码;儿童面部数据受许可限制,绝不重新分发。

英文摘要:

Facial affect models are trained almost entirely on adults, yet are increasingly applied to children in education, health, and developmental research. We present a controlled, multi-model audit of five AffectNet-pretrained expression models (EmoNet, EmotiEffLib, DDAMFN++, OpenFace 3.0, LibreFace) on children, across four child image datasets, the AffectNet-8 validation set, and two spontaneous child video datasets, through one shared harness. Three findings emerge. First, the child gap is model-agnostic: every architecture degrades from posed to naturalistic faces and shares the fear$\rightarrow$surprise confusion. Second, it is concentrated and corroborated across all five models: open-mouth faces (read as surprise, correlating with the AU26 jaw drop) and South-Asian children degrade systematically, with a smaller averted-gaze penalty, while closed-mouth faces, White and Black children, and direct gaze do not; the bias tracks expression morphology and specific populations, not skin tone. Third, the gap is diagnosable: a linear probe on frozen features reaches 0.75-0.91 on unseen children versus 0.48-0.66 zero-shot, so it lies largely in the classifier head, not the representation, whereas dimensional valence/arousal regression degrades sharply under domain shift. Building on this, recalibrating only the head on a little target data recovers $+0.13$ to $+0.28$ on the two largest child sets across all five models at negligible adult cost, though the gain is in-distribution and does not transfer across child collections. We will release the harness, per-sample predictions, and analysis code; the child face data stays license-locked and is never redistributed.

补充信息

↑