arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.19971cs.CL

基于迭代代理修正的鲁棒性不完整多模态情感分析

Robust Incomplete Multimodal Sentiment Analysis via Iterative Proxy Correction

Zhifa Geng, Subin Huang, Hao Guo, Junjie Chen, Sanmin Liu, Chao Kong

首次发表
浏览论文内容

中文总结 AI 辅助

针对不完整多模态情感分析中一次性代理初始化粗糙的问题,提出迭代代理修正框架,在MOSI等数据集上实现了优于基线的鲁棒情感预测。

中文摘要 AI 辅助

多模态情感分析旨在通过整合语言、视觉和声学线索来推断情感状态。然而,现实世界中的多模态输入常常存在不完整或损坏的情况,这会削弱跨模态互补性,并给下游融合引入误导性信息。现有的基于代理的不完整多模态情感分析(MSA)方法通常依赖一次性代理构建来补偿退化的语言信息,但生成的代理在初始化时可能较为粗糙或不可靠。过早将此类代理注入多模态推理会传播初始误差,损害情感预测。为解决这一局限,我们提出了一种用于鲁棒性不完整MSA的迭代代理修正框架。我们的方法从非语言模态构建面向语言的代理,并通过门控残差修正在多模态语境下逐步优化该代理。修正后的代理随后根据估计的语言可靠性得分与观测到的语言表示进行自适应融合,使模型能够平衡基于代理的补偿与可信的语言证据。此外,我们引入了分阶段潜在修正目标,使用完整的语言表示作为训练时的语义锚点,以稳定代理修正轨迹。在MOSI、MOSEI和SIMS数据集上针对不同缺失模态设置开展的大量实验表明,所提出的框架始终优于竞争基线,并在不完整输入下实现了鲁棒的情感预测。

英文摘要

Multimodal sentiment analysis aims to infer affective states by integrating language, visual, and acoustic cues. However, real-world multimodal inputs are often incomplete or corrupted, which can weaken cross-modal complementarity and introduce misleading information into downstream fusion. Existing proxy-based methods for incomplete MSA commonly rely on one-shot proxy construction to compensate for degraded language information, but the generated proxy may be coarse or unreliable at initialization. Prematurely injecting such a proxy into multimodal reasoning can propagate initial errors and compromise sentiment prediction. To address this limitation, we propose an iterative proxy correction framework for robust incomplete MSA. Our method constructs a language-oriented proxy from non-language modalities and progressively refines it under multimodal context through gated residual correction. The corrected proxy is then adaptively fused with the observed language representation according to an estimated language reliability score, allowing the model to balance proxy-based compensation and trustworthy linguistic evidence. In addition, we introduce a stage-wise latent correction objective that uses the complete language representation as a training-time semantic anchor to stabilize the proxy refinement trajectory. Extensive experiments on MOSI, MOSEI, and SIMS under diverse missing-modality settings demonstrate that the proposed framework consistently outperforms competitive baselines and achieves robust sentiment prediction under incomplete inputs.

发表机构

  • Renmin University of China(中国人民大学)
  • Anhui Polytechnic University(安徽工程大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑