arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

用于多任务情感行为分析的特定任务特征融合方法

Task-Specific Feature Fusion Method for Multi-Task Affective Behavior Analysis

Jiajun Sun, Zhe Gao

arXiv 2607.13986首次发表:更新:

发表机构

Shanghai Normal University(上海师范大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对ABAW11多任务情感行为分析,先适配预训练视觉骨干提取特征,再系统比较多种方法,最终选择特定任务融合和预测策略,在验证集上取得较好成绩,证明该策略简单有效。

AI 中文摘要

第11届野生情感行为分析(ABAW11)多任务学习挑战要求统一系统从官方s-Aff-Wild2图像预测效价-唤醒、分类表情和面部动作单元。尽管这些任务通过面部行为自然相关,但验证实验表明它们受益于不同视觉特征等。本文研究ABAW11多任务情感行为分析的任务自适应特征融合。先在外部面向表情的面部图像集上适配两个预训练视觉骨干并冻结以从ABAW11官方数据提取互补帧级特征,然后系统比较多种预测头、时间卷积头等。最终系统选择特定任务融合和预测策略。在ABAW11验证集上取得了一定成绩,结果表明冻结视觉特征的任务自适应融合是一种简单有效的策略。

英文摘要

The 11th Affective Behavior Analysis in-the-wild (ABAW11) Multi-Task Learning Challenge requires a unified system to predict valence-arousal, categorical expressions, and facial action units from the official s-Aff-Wild2 images. Although these tasks are naturally related through facial behavior, our validation experiments show that they benefit from different visual features, temporal processing strategies, fusion mechanisms, and calibration procedures. In this paper, we study task-adaptive feature fusion for ABAW11 multi-task affective behavior analysis. We first adapt two pretrained visual backbones, DINOv2 ViT-L and DINOv3 ConvNeXt-base, on an external expression-oriented facial image set and then freeze them to extract complementary frame-level features from the official ABAW11 data. On top of these frozen features, we systematically compare frame-level prediction heads, temporal convolutional heads, post-hoc temporal smoothing, LightGBM models, feature concatenation, gated fusion, residual fusion, late logit fusion, threshold calibration, and shared MTL structures. The final system selects task-specific fusion and prediction strategies rather than forcing all tasks to share a single architecture. On the ABAW11 validation set, the selected system achieves an EXPR macro-F1 of 0.4222, an AU macro-F1 of 0.5402, and a mean VA CCC of 0.6717, resulting in an overall validation score of 1.6341. The results suggest that task-adaptive fusion of frozen visual features is a simple and effective strategy for ABAW-style multi-task affective behavior analysis.

CommentsExtended arXiv version with an appendix. Code will be made publicly available

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑