arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.29310cs.CV

CALM-AH:面向视频层面矛盾与犹豫识别的ABAW11校准多模态集成模型,采用可靠性门控多专家共识机制

CALM-AH: An ABAW11-Calibrated Multimodal Ensemble with Reliability-Gated Multi-Expert Consensus for Video-Level Ambivalence and Hesitancy Recognition

Wenzhuo Sun, Mingjian Liang, Richard Attfield, Zongyuan Ge, Xuelian Cheng, Pamela Carreno-Medrano

首次发表
浏览论文内容

中文总结 AI 辅助

针对ABAW11视频层面矛盾与犹豫识别任务,提出CALM-AH多模态集成模型及RG-MEC门控多专家共识机制,在ABAW11数据集上分别取得0.7525与0.7771的Macro-F1性能。

中文摘要 AI 辅助

矛盾与犹豫(A/H)是可通过语言、语音、面部活动及其他非语言线索表达的微妙行为状态。ABAW11 A/H视频识别挑战赛要求系统为每段自然访谈视频分配二分类A/H标签,采用Macro-F1作为性能度量,确保A/H与非A/H样本的识别同等重要。本文提出CALM-AH,一种融合文本、声学、视觉及衍生行为统计特征的多模态集成模型,构建15种非空特征分支组合,针对每种组合通过验证集二分类交叉熵从三类分类器中选最优模型,并优化其决策阈值以提升验证集Macro-F1;所得二分类决策采用从BROTHER迁移的固定硬投票权重进行融合。本文进一步引入可靠性门控多专家共识(RG-MEC),一种锚点保留的决策级集成方法,将初始预测与三个互补修正专家(CALM-AH、AffectGPT、基于GPT的语义验证器)的预测相结合:初始系统提供默认预测,仅当三个修正专家一致支持同一替代类别时才覆盖初始标签,否则保留锚点预测。这种一致性门控设计可限制孤立专家错误的影响,同时在任务特定多模态情感证据与语义语用证据完全一致时允许双向修正。在参与者不重叠的ABAW11数据集上,CALM-AH取得0.7525的Macro-F1,完整RG-MEC系统则达到0.7771的Macro-F1。

英文摘要

Ambivalence and hesitancy (A/H) are subtle behavioural states that may be expressed through language, voice, facial activity, and other non-verbal cues. The ABAW11 A/H Video Recognition Challenge asks systems to assign a binary A/H label to each naturalistic interview video. Performance is measured using Macro-F1 so that recognition of both A/H and No-A/H samples receives equal importance. We present CALM-AH, a multimodal ensemble that combines textual, acoustic, visual, and derived behavioural-statistical features. We construct 15 non-empty combinations of these feature branches. For each combination, we select the best of three classifier families using validation binary cross-entropy and optimise its decision threshold for validation Macro-F1. The resulting binary decisions are combined using fixed hard-voting weights transferred from BROTHER. We further introduce Reliability-Gated Multi-Expert Consensus(RG-MEC), an anchor-preserving decision-level ensemble that combines an initial prediction with three complementary correction experts: CALM-AH, AffectGPT, and a GPT-based semantic verifier. The initial system provides the default prediction. Its label is overridden only when all three correction experts unanimously support the same alternative class; otherwise, the anchor prediction is retained. This unanimity-gated design limits the influence of isolated expert errors while permitting bidirectional correction when task-specific, multimodal-affective, and semantic-pragmatic evidence are fully consistent. On the participant-disjoint ABAW11 dataset, CALM-AH achieves a Macro-F1 of 0.7525, and the complete RG-MEC system achieves 0.7771.

发表机构

  • Monash University(莫纳什大学)
  • Monash Suzhou Research Institute(莫纳什苏州研究院)

机构由 AI 辅助整理,请以论文原文为准。

↑