arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.18470cs.MMcs.CL

分而治之:信息序数空间中的混合瓶颈专家用于基于视频的多模态情感分析

Divide and Conquer: Mixture-of-Bottleneck Experts in Informative Ordinal Space for Video-based Multimodal Sentiment Analysis

  • College of Computer Science and Software Engineering, Shenzhen University(深圳大学计算机科学与软件工程学院)
  • School of Electronics and Information Technology, Sun Yat-sen University(中山大学电子信息工程学院)
  • School of Cyberspace Security, Guangzhou University(广州大学网络空间安全学院)
  • School of Computer Science, South China Normal University(华南师范大学计算机学院)
  • VinUniversity(越南河内VinUniversity)
  • Nanyang Technological University(南洋理工大学)
  • School of Electrical and Electronic Engineering, Nanyang Technological University(南洋理工大学电气与电子工程学院)
  • Desay SV Automotive Co., Ltd(德赛西威汽车电子股份有限公司)

机构由 AI 辅助整理,请以论文原文为准。

Ronghao Lin, Qiaolin He, Zefeng Lu, Yichu Liu, Li Huang, Sijie Mai, Haifeng Hu, Yap-peng Tan

AI总结:

本文提出混合瓶颈专家框架,将视频多模态情感分析重构为序数回归,解耦极性与强度,通过信息瓶颈和路由融合提升性能与可解释性。

AI中文摘要:

基于视频的多模态情感分析(MSA)必须处理人类说话视频中来自文本、音频和图像序列的信息,然而当前方法往往无法以任务感知的方式整合模态。大多数模型将视频情感预测视为单一任务,忽视了其序数性质,并且它们的融合策略难以捕获跨模态的多样独特和协同线索。为解决这些局限,我们采用分而治之的视角,将MSA重新表述为序数回归问题,并将其解耦为极性识别和强度预测。受信息论驱动,我们引入了一种混合瓶颈(MoB)框架,该框架为不同模态的极性和强度特定专家分配不同的潜在表示。通过信息瓶颈的学习,每个专家学习紧凑且任务相关的表示,同时过滤掉冗余和噪声。一个多模态瓶颈路由融合模块随后利用硬挖掘策略融合这些专家潜在表示,在序数情感空间中引导预测。在4个MSA数据集和4个语言模型上的大量实验表明,MoB有效利用了来自不同模态的信息性潜在表示,并捕获了通用情感结构。除了更强的性能外,MoB还全面捕获了细粒度的模态内和模态间动态,使得对细微视频情感信号的更可信定位成为可能。

英文摘要:

Video-based Multimodal sentiment analysis (MSA) must handle information from text, audio, and image sequence in human speaking videos, yet current methods often fail to integrate modalities with task awareness. Most models treat video sentiment prediction as a single task, overlooking its ordinal nature, and their fusion strategies struggle to capture diverse unique and synergic cues across modalities. To address these limitations, we adopt a divide-and-conquer perspective by reformulating MSA as an ordinal regression problem and decoupling it into polarity recognition and intensity prediction. Driven by information theory, we introduce a Mixture-of-Bottleneck (MoB) framework that assigns different latents to polarity- and intensity-specific experts for different modalities. With the learning of information bottleneck, each expert learns compact and task-relevant representations while filtering out redundancy and noise. A multimodal bottleneck routing fusion module then fuses these expert latents with hard mining strategy, guiding the prediction in the ordinal sentiment space. Extensive experiments on 4 MSA datasets and 4 language models show that MoB effectively leverages informative latents from diverse modalities and captures general sentiment structure. Beyond stronger performance, MoB comprehensively captures fine-grained intra- and inter-modal dynamics, enabling more trustworthy localization of nuanced video sentiment signals.

↑