arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

用于情感行为分析的注意力因果监督

Causal Supervision of Attention for Affective Behaviour Analysis

Nemanja Rašajski, Konstantinos Makantasis, Antonios Liapis, Georgios N. Yannakakis

arXiv 2607.12091首次发表:更新:

发表机构

Institute of Digital Games, University of Malta; AI Department, University of Malta(马耳他大学数字游戏研究所; 马耳他大学人工智能系)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

情感行为分析面临区分情感线索与虚假因素的挑战,本文提出注意力池化框架,含因果监督、交叉协方差独立正则化及门控非线性变换三个部分,提升了泛化能力,在多任务学习挑战中取得了较好结果。

AI 中文摘要

情感行为分析旨在使机器能够在现实环境中从行为信号(特别是面部表情)推断人类情感状态。第11届野外情感行为分析竞赛包括基于s - Aff - Wild2数据库的多任务学习挑战,参与者要为效价 - 唤醒估计、表情识别和动作单元检测开发统一框架。这具有挑战性,因为要区分与情感相关的线索和诸如身份、光照、姿势及人口统计学差异等虚假因素。注意力机制虽能聚合信息,但可能利用特定数据集的相关性而非真实情感线索。为提高泛化能力,我们提出注意力池化框架,它能促进主体不变的注意力,同时增加特征表现力。我们的方法由三个部分组成:引入因果监督以强制关注跨主体具有不变预测值的面部区域;对键(K)和值(V)投影应用交叉协方差独立正则化以鼓励互补、非冗余表示;用门控非线性SwiGLU变换替换线性值投影以增加特征表现力并捕捉更细粒度的情感线索。我们的方法在官方验证集上,效价 - 唤醒估计的CCC_VA = 0.5123,表情识别的F1_EX = 0.3116,动作单元检测的F1_AU = 0.3974,总体P分数为1.2214。

英文摘要

The \textit{11th Affective Behaviour Analysis in-the-wild Competition} includes the Multi-Task Learning Challenge, where participants develop a unified framework for Valence-Arousal Estimation, Expression Recognition, and Action Unit Detection. The challenge lies in learning emotion-related representations that generalize across subjects while remaining robust to spurious factors such as identity, illumination, pose, and demographic variation. To aggregate features extracted by a pre-trained backbone into a compact representation for prediction, attention mechanisms selectively weight the most informative facial regions. However, these attention weights can still capture dataset-specific correlations rather than genuine affective cues. To address this limitation, we propose an attention pooling framework that combines causal supervision with cross-covariance regularization of attention components, encouraging subject-invariant attention and non-redundant representations that improve generalization. Our method achieves $CCC_{VA}=0.5123$ for VA estimation on the official validation set, together with $F_{EX}=0.3116$ and $F_{AU}=0.3974$ for expression recognition and action unit detection, respectively, resulting in an overall $P$ score (the sum of the individual task metrics) of $1.2214$.

Comments10 pages, 1 figure, 2 tables

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑