视频中矛盾与犹豫识别的简单特征与诚实校准
Simple Features and Honest Calibration for Ambivalence and Hesitancy Recognition in Video
- Indian Institute of Science Education and Research Bhopal(印度科学教育与研究学院博帕尔分院)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
针对ABAW 2026 BAH挑战赛中的矛盾与犹豫识别问题,系统结合情感专用多模态表示与语言犹豫线索,经AMF融合及AP加权集成。引入‘ASR擦除时间’构建特征,实验表明语言是最强通道,校准比架构重要,固定阈值AP加权测试集得分0.731。
AI中文摘要:
我们研究了ABAW 2026 BAH挑战赛中的矛盾与犹豫(A/H)识别问题:给定一段简短采访视频,预测人物是否表现出A/H迹象。我们的系统将情感专用的文本、音频和视觉表示与一小组可读的语言犹豫线索相结合,通过我们称为情感标记融合(AMF)的可靠性门进行融合,并在固定决策阈值下通过简单的AP加权集成完成。我们还引入了‘ASR擦除时间’,基于此构建的16个特征形成了最强且最独立的非语言通道。通过控制实验发现:跨模态冲突设计对BAH帮助不大;语言是最强通道,情感专用音频是有用的第二通道;校准比架构更重要。固定阈值下的AP加权在测试集上达到0.731。
英文摘要:
We address ambivalence and hesitancy (A/H) recognition in the ABAW 2026 BAH Challenge: given a short interview video, predict whether the person shows signs of A/H. Our system combines affect-specialised text, audio, and visual representations with a small set of readable linguistic hesitation cues, fused by a reliability gate we call Affective Marker Fusion (AMF), and finished with a simple AP-weighted ensemble at a fixed decision threshold. We also introduce \emph{ASR-erased time}: speech recognisers delete fillers and hesitation pauses from the transcript, but the chunk timestamps keep the time those events took, and sixteen features built from these gaps form the strongest and most independent non-verbal channel we measured (AP $0.718$, correlation $0.11$--$0.36$ with all other members). Across controlled experiments we find three things: cross-modal conflict design does not reliably help on BAH; language is by far the strongest channel while affect-specialised audio is a useful second; and calibration matters more than architecture. Fitting ensemble weights and a threshold on the small validation split overfits: it scores $0.741$ macro-F1 on validation but only $0.690$ on the untouched test set. AP-weighting at a fixed threshold instead reaches $\mathbf{0.731}$ on test.