分离、细化、整合:用于情感分析的多模态融合分解
Segregate, Refine, Integrate: Decomposing Multimodal Fusion for Sentiment Analysis
- National Technical University of Athens(雅典国立技术大学)
- Athena Research Center(雅典娜研究中心)
- University of Bern(伯尔尼大学)
- Archimedes AI(阿基米德人工智能公司)
- Synaptic Bloom PBC(突触绽放有限责任公司)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
研究情感分析中的多模态融合问题,提出SeRIn方案,通过分离特定模态与跨模态路径,各自细化演变,将跨模态交互推迟到预测步骤,在CH - SIMS和CMU - MOSEI上取得领先成果。
AI中文摘要:
多模态融合必须同时细化特定模态信号并对跨模态交互进行建模;这两个相互竞争的目标通常在同一操作中纠缠在一起。我们提出了SeRIn(分离、细化、整合),一种多模态语言模型融合方案,将这种分离作为一种架构先验。特定模态表示沿着孤立路径演变,各自根据其编码器上下文进行细化,而专用的跨模态路径累积它们的联合演变而不污染单模态流。完整的跨模态交互推迟到最终预测步骤——消融实验证实结构化交互而非增加的容量推动了性能提升;视觉损坏下的门控分析揭示了在没有明确监督的情况下出现的模态重新加权。SeRIn在CH - SIMS和CMU - MOSEI上取得了领先成果,改进了两个基准测试的所有指标。
英文摘要:
Multimodal fusion must simultaneously refine modality-specific signals and model cross-modal interactions; two competing objectives typically entangled within the same operation. We propose \textbf{SeRIn} (\textbf{Se}gregate, \textbf{R}efine, \textbf{In}tegrate), a multimodal LM fusion scheme that enforces this separation as an architectural prior. Modality-specific representations evolve along isolated pathways, each refined against its respective encoder context, while a dedicated cross-modal pathway accumulates their joint evolution without contaminating unimodal streams. Full cross-modal interaction is deferred to a final prediction step - ablations confirm that structured interactions, not added capacity, drive the gains; gate analysis under visual corruption reveals emergent modality reweighting without explicit supervision. SeRIn achieves state-of-the-art results on CH-SIMS and CMU-MOSEI, improving all metrics on both benchmarks.