arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.31399cs.LGeess.SP

差分注意力解锁互补的脑电与语音融合用于情绪识别

Differential Attention Unlocks Complementary EEG and Speech Fusion for Emotion Recognition

Philip H. Lee, Shreeram Suresh Chandra, John H. L. Hansen

首次发表
浏览论文内容

中文总结 AI 辅助

针对多模态情绪识别中脑电噪声破坏融合的问题,提出EmoSpeechBrain框架,利用差分注意力消除脑电噪声并结合门控适配器融合脑电与语音,在PME4和EAV数据集上显著提升准确率。

中文摘要 AI 辅助

多模态情绪识别(MER)越来越多地将脑电(EEG)与语音配对,将内部神经信号和外部语音表达视为情感的信息性视图。在实践中,朴素融合的性能不如较强的单一模态,因为脑电伪影注入噪声,破坏了共享表示。我们提出了EmoSpeechBrain,一个基于噪声抑制是有效融合的前提这一见解的多模态框架。其脑电编码器使用差分注意力,取两个注意力图之间的差值来消除共享噪声并分离出判别性神经活动。一个基于注意力的门控适配器将两种模态在共享空间中对齐,并加权各自对预测的贡献。在两个数据集——PME4和EAV上,EmoSpeechBrain相较于其他最先进的(SOTA)脑电编码器将MER准确率提高了最多12.9%,并分别超过单模态语音和脑电基线最多13.1%和23.1%。这些结果表明,一旦脑电噪声被抑制,融合带来的增益是朴素组合无法实现的。

英文摘要

Multimodal emotion recognition (MER) increasingly pairs EEG with speech, treating internal neural signals and external vocal expression as informative views of affect. In practice, naive fusion underperforms the stronger single modality, because EEG artifacts inject noise that corrupts the shared representation. We introduce EmoSpeechBrain, a multimodal framework built on the insight that noise suppression is a precondition for effective fusion. Its EEG encoder uses differential attention, taking the difference between two attention maps to cancel shared noise and isolate discriminative neural activity. An attention-based gating adapter aligns both modalities in a shared space and weights each one's contribution to the prediction. On two datasets - PME4 and EAV, EmoSpeechBrain improves MER accuracy by up to 12.9% over other state-of-the-art (SOTA) EEG encoders, and surpasses unimodal speech and EEG baselines by up to 13.1% and 23.1%. These results show that once EEG noise is suppressed, fusion delivers gains that naive combination cannot.

发表机构

  • Johns Hopkins University(约翰霍普金斯大学)
  • University of Texas at Dallas(德克萨斯大学达拉斯分校)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑