arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.13857cs.LG

ReH-FUSE:面向对话中多模态情感识别的可靠性感知专家层级融合

ReH-FUSE: Reliability-Aware Hierarchical Fusion of Experts for Multimodal Emotion Recognition in Conversation

  • National Taiwan University of Science and Technology(国立台湾科技大学)

机构由 AI 辅助整理,请以论文原文为准。

Guan-Hua Wen, Hou-Chiang Tseng, Kuan-Yu Chen

AI总结:

提出ReH-FUSE,一种可靠性感知的层级专家融合框架,通过决策级路由器平衡单模态与跨模态信息,在IEMOCAP和MELD上显著提升多模态情感识别性能。

AI中文摘要:

对话中的多模态情感识别(ERC)需要适应不同证据来源的实例依赖可靠性。词汇内容可能具有决定性,语音表达可能提供补充线索,或准确识别可能需要跨模态交互;固定融合并未明确考虑这种变化。我们提出ReH-FUSE,一种具有对话感知文本、音频和跨模态专家的可靠性感知框架。其决策级路由器首先建模文本与音频之间的相对偏好,然后将所得单模态混合与跨模态专家进行平衡。这种分解将单模态竞争与跨模态选择分离开来。在IEMOCAP上的三次独立运行中,ReH-FUSE达到74.34%的加权F1和73.11%的宏F1;在MELD上达到68.03%的加权F1。受控消融实验表明,学习路由优于均匀专家平均,并受益于跨模态交互。

英文摘要:

Multimodal emotion recognition in conversation (ERC) requires adapting to the instance-dependent reliability of different evidence sources. Lexical content may be decisive, vocal expression may provide complementary cues, or accurate recognition may require cross-modal interaction; fixed fusion does not explicitly account for this variation. We propose ReH-FUSE, a reliability-aware framework with dialogue-aware text, audio, and cross-modal experts. Its decision-level router first models the relative preference between text and audio and then balances the resulting unimodal mixture against the cross-modal expert. This factorization separates unimodal competition from cross-modal selection. Across three independent runs on IEMOCAP, ReH-FUSE achieves 74.34% weighted F1 and 73.11% macro F1; on MELD, it achieves 68.03% weighted F1. Controlled ablations show that learned routing outperforms uniform expert averaging and benefits from cross-modal interaction.

补充信息

↑