arXivDaily arXiv每日学术速递 周一至周五更新
arXiv 2607.13976cs.CV

CF-Net:用于矛盾/犹豫识别的具有说话者归一化和确定性加权的冲突融合

CF-Net: Conflict Fusion with Speaker Normalisation and Certainty Weighting for Ambivalence/Hesitancy Recognition

Tung Hung Bui, Hong Hai Nguyen, Van Thong Huynh

AI总结:

针对无约束视频中矛盾和犹豫检测的挑战,提出CF-Net,用特定主干编码多模态流,经归一化和冲突融合模块处理,结合多种训练方法,在相关数据集上取得了较好成绩。

AI中文摘要:

在无约束视频中检测矛盾和犹豫具有挑战性,因为目标信号本质上模糊且通过微妙的跨模态不一致来表达。我们提出了CF-Net,这是一个提交给第三届矛盾/犹豫视频识别挑战赛(ABAW 第11届,ECCV 2026)针对BAH数据集的深度多模态网络。CF-Net用冻结的SigLIP2、HuBERT和DistilBERT主干对视觉、音频和转录流进行编码,对每个说话者的主干特征进行归一化以减少身份泄露,并通过冲突融合模块进行融合。训练结合了确定性加权焦点损失、流形混合和模态丢弃。CF-Net在BAH验证集上Macro F1为0.7155,在私人挑战测试集上为0.7364(AP = 0.7492)。

英文摘要:

Detecting ambivalence and hesitancy (AH) in unconstrained video is challenging because the target signal is inherently ambiguous and expressed through subtle cross-modal incongruence rather than prototypical affect. We present CF-Net, a deep multimodal network submitted to the 3rd Edition of the AH Video Recognition Challenge (ABAW 11th, ECCV 2026), targeting the BAH dataset. CF-Net encodes visual, audio, and transcript streams with frozen SigLIP2, HuBERT, and DistilBERT backbones, normalises backbone features per speaker to reduce identity leakage, and fuses them via a ConflictFusion module that explicitly computes pairwise cross-modal incongruence. Training combines certainty-weighted focal loss, manifold mixup, and modality dropout; an auxiliary certainty-regression head uses ambiguity annotations to stabilise learning on genuinely borderline samples. CF-Net achieves a Macro F1 of 0.7155 on the BAH validation set and 0.7364 (AP = 0.7439) on the private challenge test set.

补充信息

↑