第11届ABAW挑战赛中的HSEmotion团队:多任务学习与矛盾/犹豫视频识别
HSEmotion Team at the 11th ABAW Challenge: Multi-Task Learning and Ambivalence/Hesitancy Video Recognition
浏览论文内容
中文总结 AI 辅助
在第11届ABAW挑战赛中,团队用冻结轻量级面部提取器等方法进行多任务学习,在s - Aff - Wild2数据集上超ConvNeXt基线;扩展视听管道用于矛盾/犹豫视频识别,在BAH数据集上提升性能,无需微调重型骨干,展示了轻量级融合的优势。
中文摘要 AI 辅助
本文展示了我们在第11届自然环境下情感行为分析(ABAW)竞赛中的成果。对于在s - Aff - Wild2数据集上同时预测效价、唤醒度、面部表情和动作单元的多任务学习,我们使用冻结的轻量级面部提取器MT - EmotiDDAMFN和MT - EmotiEffNet - B0,并采用单独的头部和系统的后处理方法,包括时间高斯平滑、每类表情偏差、AffectNet融合、每个动作单元的阈值调整以及加权骨干融合。在官方验证集上,我们的集成显著超过了ConvNeXt基线的性能。对于在扩展的BAH数据集上的矛盾/犹豫视频识别,我们通过面部、HuBERT音频和RoBERTa文本分类器的后期融合、时间聚合和全局文本门,将视听管道扩展到视频级Macro F1。验证集上的帧级加权F1从ABAW - 8中的0.74提高到0.79,而最佳公开测试视频级Macro F1达到0.73。在这两个任务中,无需微调重型骨干就能取得有竞争力的性能。这些结果表明,系统的预测校准和轻量级多模态融合可以与实质上更重的端到端方法相媲美,同时提高了效率和部署灵活性。
英文摘要
This article presents our results for the 11th Affective Behavior Analysis in-the-Wild (ABAW) competition. For multi-task learning with simultaneous prediction of valence, arousal, facial expressions, and action units on s-Aff-Wild2 dataset, we use frozen lightweight facial extractors, MT-EmotiDDAMFN and MT-EmotiEffNet-B0, with separate heads and systematic post-processing: temporal Gaussian smoothing, per-class expression bias, AffectNet blending, per-AU threshold tuning, and weighted backbone fusion. On the official validation set, our ensemble significantly exceeds the performance of the ConvNeXt baseline. For ambivalence/hesitancy video recognition on the expanded BAH dataset, we extend the audiovisual pipeline to video-level Macro F1 by late fusion of face, HuBERT audio, and RoBERTa text classifiers, temporal aggregation, and a global-text gate. Frame-level Weighted F1 on validation set rises from 0.74 in ABAW-8 to 0.79, while the best public-test video-level Macro F1 reaches 0.73. In both tasks, competitive performance is achieved without fine-tuning heavy backbones. These results indicate that systematic prediction calibration and lightweight multimodal fusion can rival substantially heavier end-to-end approaches while offering improved efficiency and deployment flexibility.
发表机构
- Central University(中央大学)
- Sber AI Lab(Sber人工智能实验室)
- HSE University(俄罗斯高等经济研究大学)
机构由 AI 辅助整理,请以论文原文为准。