发表机构
National University of Defense Technology(国防科技大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出BC-MLF,通过分支校准任务头和融合令牌对比学习显式建模预测层分支分配,在不改融合主干下提升多模态情感分析性能,在CMU-MOSEI和CH-SIMS上取得最优结果。
AI 中文摘要
多模态情感分析整合了文本、声学和视觉线索,然而当前基于语言模型的融合方法通常隐式地处理预测层分支分配。我们引入了分支校准多模态语言融合(BC-MLF),通过分支校准任务头(BCHead)显式建模预测层分支分配,并辅以融合令牌对比学习(FTCL)进行情感感知的融合令牌正则化。FTCL根据连续情感亲和度组织平均池化的融合令牌表示,而BCHead通过轻量级样本自适应约束混合,结合融合、文本和音视频预测。在不修改融合主干的情况下,BC-MLF一致地改进了复现的DeepMLF基线,并在CMU-MOSEI和CH-SIMS上,在分类和回归指标上取得了比较方法中最强的结果。受控消融实验表明,样本自适应的预测层分支分配始终优于静态分支聚合。代码可在该https URL获取。
英文摘要
Multimodal sentiment analysis integrates textual, acoustic and visual cues, yet current language-model-based fusion methods typically leave prediction-layer branch allocation implicit. We introduce Branch-Calibrated Multimodal Language Fusion (BC-MLF), which explicitly models prediction-layer branch allocation through a Branch-Calibrated Task Head (BCHead), complemented by Fusion Token Contrastive Learning (FTCL) for sentiment-aware fusion-token regularization. FTCL organizes mean-pooled fusion-token representations according to continuous sentiment affinity, while BCHead combines fusion, text and audiovisual predictions through a lightweight sample-adaptive constrained mixture. Without modifying the fusion backbone, BC-MLF consistently improves the reproduced DeepMLF baseline and achieves the strongest results among the compared methods on CMU-MOSEI and CH-SIMS across classification and regression metrics. The controlled ablations show that sample-adaptive prediction-layer branch allocation consistently outperforms static branch aggregation. Code is available at https://github.com/sunyulin0421/BC-MLF.
Comments5 pages, 3 figures, 3 tables