发表机构
University of Paris 8; LIASD Laboratory(巴黎第八大学; LIASD实验室)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
为ECCV 2026的ABAW11挑战赛开发多模态框架,用预训练编码器提取多模态特征,建立单模态基线,在此基础上提出含交叉注意力和门控融合的多模态架构,验证集宏F1达0.7394,提升显著。
AI 中文摘要
我们提出了一个用于视频中矛盾/犹豫(A/H)识别的多模态框架,该框架是为2026年ECCV的ABAW11挑战赛而开发的。所提出的方法使用三个预训练编码器融合从BAH数据集中提取的文本、声学和视觉模态:用于转录本的F2LLM-v2-0.6B(1024维)、用于音频的WavLM-Large(1024维)和用于面部视频的VideoMAE V2(768维)。我们首先使用经典分类器(MLP、随机森林、梯度提升决策树)建立全面的单模态基线,每个分类器通过Optuna进行优化,并仅使用文本特征在测试集上获得了最佳单模态宏F1为0.6659,显著优于零样本Video-LLaVA基线(宏F1:0.2827)。在此基础上,我们提出了一种多模态融合架构,该架构将所有三种模态的双向交叉注意力与门控多模态单元(GMU)相结合,架构和优化超参数均通过50次试验的Optuna搜索选择。该模型在验证集上实现了宏F1为0.7394,相对于最佳单模态基线有11.0%的相对提升,证实了显式跨模态交互捕获了单个模态无法单独提供的互补线索。使用该模型在官方未标记的私人测试集上生成最终预测,并根据挑战协议提交。代码可在这个https URL上公开获取。
英文摘要
We present a multimodal framework for Ambivalence/Hesitancy (A/H) recognition in video, developed for the ABAW11 challenge at ECCV 2026. The proposed approach fuses textual, acoustic, and visual modalities extracted from the BAH dataset using three pretrained encoders: F2LLM-v2-0.6B for transcripts (1024-d), WavLM-Large for audio (1024-d), and VideoMAE V2 for facial video (768-d). We first establish comprehensive unimodal baselines using classical classifiers (MLP, Random Forest, GBDT), each optimized via Optuna, and obtain a best unimodal Macro F1 of \textbf{0.6659} on the test set using text features alone -- substantially outperforming the zero-shot Video-LLaVA baseline (Macro F1: 0.2827). Building on these baselines, we propose a multimodal fusion architecture that combines bidirectional cross-attention across all three modalities with a Gated Multimodal Unit (GMU), with both architectural and optimization hyperparameters selected through a 50-trial Optuna search. This model achieves a Macro F1 of \textbf{0.7394} on the validation set, a relative improvement of 11.0\% over the best unimodal baseline, confirming that explicit cross-modal interaction captures complementary cues that no single modality provides in isolation. Final predictions on the official, unlabeled private test set are generated using this model and submitted according to the challenge protocol. Code is publicly available at https://github.com/yassineouzar/IUSD_AH/