发表机构
Jianghan University(江汉大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究针对多模态情感分析的静态计算图与忽略情感序数层级的问题,提出MAESTRO框架,通过文本引导的MoE机制与O-PCL实现动态多模态表征,在CMU-MOSI等基准上取得最优性能。
AI 中文摘要
多模态情感分析(Multimodal Sentiment Analysis, MSA)是情感计算的核心组成部分,旨在通过整合语言内容与语音语调、面部微表情等非语言线索,解读复杂的情感状态。尽管近期基于解耦的方法推动了该领域发展,但仍受限于两大方法学挑战:其一,静态计算图无论语义复杂度如何均不加区分地处理所有样本,导致对不同情感表达与上下文场景的表征效果欠佳;其二,通用对比目标常忽略情感强度的内在序数层级。为系统解决这些局限,我们提出结合文本路由与序数原型优化的多模态自适应专家选择框架(Multimodal Adaptive Expert Selection with Text Routing and Ordinal Prototype Optimization, MAESTRO),该框架旨在动态协调与优化多模态表征。受管弦乐队指挥启发,我们设计了文本引导的混合专家混合(Mixture-of-Experts, MoE)机制;与静态融合不同,该模块利用语言上下文作为路由信号,动态激活特定的视听专家,从而通过自适应特征增强解决跨模态歧义。此外,为捕捉细粒度的情感梯度,我们提出序数感知原型对比学习(Ordinal-aware Prototype Contrastive Learning, O-PCL);通过在原型学习目标中引入基于距离的惩罚项,O-PCL构建了保留情感自然顺序的结构化潜在空间。在CMU-MOSI与CMU-MOSEI基准上开展的大量实验表明,MAESTRO实现了当前最优性能,定性分析进一步证实了我们的动态路由范式的可解释性。
英文摘要
Multimodal Sentiment Analysis (MSA) is a fundamental component of affective computing that aims to decipher complex emotional states by integrating verbal content with non-verbal cues including vocal intonation and facial micro-expressions. While recent disentanglement-based approaches have advanced the field, their potential is hindered by two methodological challenges. First, static computation graphs process all samples indiscriminately regardless of semantic complexity, which leads to suboptimal representation for diverse emotional expressions and contextual scenarios. Second, generic contrastive objectives often neglect the intrinsic ordinal hierarchy of sentiment intensities. To systematically address these limitations, we introduce Multimodal Adaptive Expert Selection with Text Routing and Ordinal prototype optimization (MAESTRO), a novel framework designed to dynamically orchestrate and refine multimodal representations. Drawing inspiration from an orchestra conductor, we design a Text-Guided Hybrid Mixture-of-Experts (MoE) mechanism. Unlike static fusion, this module utilizes linguistic context as a routing signal to dynamically activate specific audio-visual experts, thereby resolving cross-modal ambiguity through adaptive feature enhancement. Furthermore, to capture fine-grained sentiment gradations, we propose an Ordinal-aware Prototype Contrastive Learning (O-PCL). By incorporating distance-based penalties into the prototype learning objective, O-PCL enforces a structured latent space that preserves the natural order of emotion. Extensive experiments on the CMU-MOSI and CMU-MOSEI benchmarks demonstrate that MAESTRO achieves state-of-the-art performance, and qualitative analysis further confirms the interpretability of our dynamic routing paradigm.