发表机构
St. Petersburg Federal Research Center of the Russian Academy of Sciences (SPC RAS); HSE University; ITMO University(俄罗斯科学院圣彼得堡联邦研究中心; 圣彼得堡高等经济学院; 圣彼得堡国立信息技术机械与光学大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对第十一届ABAW竞赛的视频级矛盾情绪和犹豫情绪识别,提出以文本为中心的多模态方法,用文本残差融合模型结合多种特征,实验表明该方法能有效提高识别性能,无需大型模型集成。
AI 中文摘要
自动识别矛盾情绪和犹豫情绪具有挑战性,因为这些状态可能通过不一致的语言、声学、面部和情境模式来表达,而表现最佳的系统通常依赖于计算成本高昂的集成方法。我们为第十一届野外情感与行为分析(ABAW)挑战赛提出了一种以文本为中心的视频级矛盾情绪和犹豫情绪识别的单模态方法。该方法使用以文本为中心的多模态融合模型结合语言、声学、面部和场景特征。文本残差融合将文本视为锚定模态,并根据其他模态进行门控残差调整。在行为矛盾/犹豫(BAH)语料库上的实验证实文本是最强的单模态。文本残差融合模型在开发集和公共测试子集上的平均宏F1分数(MF1)为75.14%。在私人测试子集上,它达到了78.24%的MF1,比文本模型高出4.03%。这些结果表明互补的多模态信息可以提高识别性能,而无需大型模型集成。
英文摘要
Automatic recognition of ambivalence and hesitancy is challenging because these states may be expressed through inconsistent linguistic, acoustic, facial, and contextual patterns, while top-performing systems often rely on computationally expensive ensembles. We present a single text-centered multimodal approach for video-level ambivalence and hesitancy recognition for the 11th Affective & Behavior Analysis in-the-Wild (ABAW) Challenge. The proposed approach combines linguistic, acoustic, facial, and scene features using text-centered multimodal fusion model. Text Residual Fusion treats text as the anchor modality and applies gated residual adjustments based on the other modalities. Experiments on the Behavioural Ambivalence/Hesitancy (BAH) corpus confirm that text is the strongest unimodal modality. The Text Residual Fusion model achieves an average Macro F1-score (MF1) of 75.14% across the Development and Public Test subsets. On the Private Test subset, it reaches an MF1 of 78.24%, outperforming the text model by 4.03%. These results demonstrate that complementary multimodal information can improve recognition performance without requiring a large model ensemble.
Comments10 pages, 2 figures