发表机构
University of the Basque Country (UPV/EHU); Department of Innovation Engineering, University of Salento & Institute of Applied Sciences and Intelligent Systems (CNR); Faculty of Data Science and Computing, Universiti Malaysia Kelantan(巴斯克大学(UPV/EHU); 萨伦托大学创新工程系与应用科学和智能系统研究所(CNR); 马来西亚吉兰丹大学数据科学与计算学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究针对第11届ABAW挑战赛,提出AffectFlow-DINO多任务学习系统,通过条件整流流头建模面部行为模糊性,经蒙特卡洛采样实现不确定性感知预测,结合多种策略提升性能,最终大幅超越官方挑战赛基线。
AI 中文摘要
我们提出了AffectFlow-DINO,这是一个用于第11届ABAW挑战赛的多任务学习系统,它通过条件整流流头扩展了标准确定性架构,以对野生面部行为的内在模糊性进行建模。该模型学习条件生成分布,通过蒙特卡洛采样实现不确定性感知的一对多预测。系统能联合估计连续的效价-唤醒值、分类八种面部表情并从静态面部图像中检测十二个动作单元。基于冻结的DINOv3 ViT-S/16骨干网络,大量消融研究表明整流流解码持续改进确定性预测,特别是效价-唤醒估计(CCC-V提高0.058)。我们还表明事后阈值校准能在不重新训练的情况下有效恢复严重不平衡罕见类别的性能(如恐惧类别从3.8%提升到33.1%)。结合骨干网络微调与流调整,最终模型实现了P_MTL = 1.177,大幅超越官方挑战赛基线P_MTL = 0.45。
英文摘要
We present \textbf{AffectFlow-DINO}, a multi-task learning system for the 11th ABAW challenge that extends a standard deterministic architecture with a conditional rectified-flow head to model the inherent ambiguity of in-the-wild facial behavior. Instead of predicting a single affect estimate, the model learns a conditional generative distribution, enabling uncertainty-aware one-to-many predictions through Monte Carlo sampling. The system jointly estimates continuous valence-arousal, classifies eight facial expressions, and detects twelve Action Units from static face images. Built on a frozen DINOv3 ViT-S/16 backbone, extensive ablation studies show that rectified-flow decoding consistently improves deterministic prediction, particularly for valence-arousal estimation (CCC-V $+0.058$). We further show that post-hoc threshold calibration effectively recovers performance on severely imbalanced rare classes (e.g., Fear: $3.8\% \rightarrow 33.1\%$) without retraining. Combined with backbone fine-tuning and flow retuning, the final model achieves $\mathbf{P_{MTL}=1.177}$, substantially outperforming the official challenge baseline of $P_{MTL}=0.45$.