arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.17130cs.CV

预测人类分歧以实现校准的动态面部表情识别

Predicting Human Disagreement for Calibrated Dynamic Facial Expression Recognition

Yiming Wang, Frederick W. B. Li, Jingyun Wang

首次发表
浏览论文内容

中文总结 AI 辅助

提出分歧感知的DFER框架,用Dirichlet-Multinomial似然直接训练标注计数,结合模糊性头与拒绝规则,在DFEW上降ECE 30%、AURC 15%,并迁移至FERV39k。

中文摘要 AI 辅助

动态面部表情识别(DFER)基准如DFEW为每个视频片段提供多个标注者的投票,但大多数模型将其压缩为多数标签,无法在推理时表示人类分歧。我们提出了一种分歧感知的DFER框架,该框架直接基于原始标注者计数向量进行训练,使用Dirichlet-Multinomial似然。与仅均值的软标签目标不同,所提出的似然为Dirichlet浓度提供了尺度敏感的监督,同时保留了预测均值。一个独立的模糊性头预测未见片段的标注熵,一个单调的Chow式拒绝规则结合预测模糊性、空虚性、时间不稳定性和输入质量进行选择性预测。在DFEW上,该方法在保持识别精度的同时,将ECE降低了30%,AURC降低了15%,预测模糊性与测试片段的标注熵达到0.52的Spearman相关性。校准和选择性预测的收益可迁移至FERV39k,并在身份和电影不相交的DFEW划分下保持。

英文摘要

Dynamic facial expression recognition (DFER) benchmarks such as DFEW provide multiple annotator votes per clip, yet most models collapse them to a majority label and cannot represent human disagreement at inference time. We propose a disagreement-aware DFER framework that trains directly on the raw annotator count vector using a Dirichlet-Multinomial likelihood. Unlike mean-only soft-label objectives, the proposed likelihood provides scale-sensitive supervision for the Dirichlet concentration while preserving the predictive mean. A separate ambiguity head predicts annotation entropy for unseen clips, and a monotone Chow-style reject rule combines predicted ambiguity, vacuity, temporal instability, and input quality for selective prediction. On DFEW, the method preserves recognition accuracy while reducing ECE by 30% and AURC by 15%, and predicted ambiguity reaches a Spearman correlation of 0.52 with the annotation entropy of test clips. The calibration and selective-prediction gains transfer to FERV39k and remain under identity- and movie-disjoint DFEW splits.

发表机构

  • Durham University(杜伦大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑