用于部分标记多任务面部表情识别的共享潜在空间
A Shared Latent for Partially-Labeled Multi-Task Facial Affect Recognition
- Faculty of Computer Science and Engineering, Ho Chi Minh City University of Technology (HCMUT)(胡志明市理工大学计算机科学与工程学院)
- Dept. of AI, FPT University(FPT大学人工智能系)
- Department of Artificial Intelligence Convergence, Chonnam National University(全南国立大学人工智能融合系)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
研究针对野外面部表情多任务且标注不完整不均衡的问题,提出将部分标记多任务学习视为对共享情感潜在空间的边缘化方法,在s-Aff-Wild2数据集上提升了表情和动作单元识别效果,得出相关结论并展示了小范围转移结果。
AI中文摘要:
野外面部表情本质上是多任务的:效价-唤醒、离散表情和面部动作单元描述同一张脸。但真实语料库对这些任务的标注不完整且不均衡,多数系统掩盖缺失标签或插入伪标签,忽略跨任务信号。我们将部分标记多任务学习视为对共享情感潜在空间的边缘化:一个变分瓶颈介导所有三个任务解码器,一个任务标注的帧会塑造其他任务所用的表示,掩码目标作为证据下界的重建项重新出现。在s-Aff-Wild2数据集上,该方法提升了表情宏F1值,打破了动作单元的上限,效价-唤醒保持在噪声范围内。每个增益都由匹配控制负项约束,表明罕见类失败是代表性的,而非损失塑造问题。我们报告了组合多任务分数,并基于控制比较得出结论,还展示了向AffectNet和RAF-DB的小范围、依赖于机制的表情优势转移作为探索性而非确定性结果。
英文摘要:
Facial affect in the wild is naturally multi-task: valence-arousal, discrete expressions, and facial action units describe the same face. Yet real corpora annotate these tasks only partially and unevenly, so most systems mask the missing labels or impute pseudo-labels and forgo the cross-task signal. We instead cast partially-labeled multi-task learning as marginalization over a shared affect latent: one variational bottleneck mediates all three task decoders, so a frame annotated for one task shapes the representation the others use, and the masked objective reappears as the reconstruction term of an evidence lower bound. On s-Aff-Wild2, where only 37% of frames carry all three labels, the classes are severely imbalanced, and pretraining on the source data is disallowed, we isolate where this coupling acts. On a single backbone it lifts expression macro-F1 from 0.403 for a dedicated specialist to 0.446, which the masked-loss model does not reach; a second, near-peer backbone with decorrelated errors then breaks an action-unit ceiling that external action-unit data could not, while valence-arousal stays within noise. Every gain is disciplined by a matched-control negative; together these controls indicate that the rare-class failure is representational, not a matter of loss shaping. As each task's source is chosen on the evaluation split, we report the assembled result, a combined multi-task score of 1.679 on validation, as an in-sample endpoint and rest our conclusions on the controlled comparisons; a small, regime-dependent transfer of the expression advantage to AffectNet and RAF-DB is presented as exploratory rather than conclusive.