发表机构
Industrial University of HCM City; Le Hong Phong High School for the Gifted; FPT University; University of Science, VNU-HCM(胡志明市工业大学; 黎鸿峰天才高中; 菲普特大学; 胡志明市国家大学科学大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
提出轻量级多模态情感识别框架RAFM_SER++,采用非对称残差注意力融合机制,在IEMOCAP和ESD上以更少参数和更快推理超越基线,适用于实时监控。
AI 中文摘要
近年来,多模态语音情感识别(SER)系统通过交互密集的跨模态变换器实现了高精度,但其计算成本限制了在延迟敏感和资源受限的监控系统中的部署。为解决这一挑战,我们提出了RAFM_SER++,一种轻量级多模态SER框架,其特点是采用非对称残差注意力融合机制(RAFM)。RAFM不依赖计算成本高昂的双向交互,而是通过单向残差注意力路径将情感语音线索注入语义文本表示中。结合受BYOL启发的跨模态对齐目标和注意力引导池化,所提框架在保持低计算开销的同时提升了多模态表示学习。在IEMOCAP和ESD基准上的实验表明,RAFM_SER++始终优于HuBERT-Base基线,并与最先进的MemoCMT相比实现了更优的精度-效率权衡。具体而言,RAFM_SER++将可训练参数减少了60%以上,实现了更快的推理速度(79.60 it/s),并在IEMOCAP上达到81.10%的BACC分数,在ESD上达到95.39%的BACC分数。这些结果表明,轻量级非对称多模态融合是实时监控应用中交互密集跨模态变换器的有效替代方案。
英文摘要
Recent multimodal Speech Emotion Recognition (SER) systems achieve high accuracy through interaction-heavy cross-modal transformers, but their computational cost limits deployment in latency-sensitive and resource-constrained surveillance systems. To address this challenge, we propose RAFM_SER++, a lightweight multimodal SER framework featuring an asymmetric Residual Attention Fusion Mechanism (RAFM). Rather than relying on computationally expensive bidirectional interactions, RAFM injects affective speech cues into semantic text representations through a one-directional residual attention pathway. Combined with a BYOL-inspired cross-modal alignment objective and attention-guided pooling, the proposed framework improves multimodal representation learning while maintaining low computational overhead. Experiments on the IEMOCAP and ESD benchmarks demonstrate that RAFM_SER++ consistently outperforms the HuBERT-Base baseline and achieves a superior accuracy-efficiency trade-off compared with the state-of-the-art MemoCMT. Specifically, RAFM_SER++ reduces trainable parameters by more than 60%, achieves faster inference (79.60 it/s), and attains BACC scores of 81.10% on IEMOCAP and 95.39% on ESD. These results indicate that lightweight asymmetric multimodal fusion is an effective alternative to interaction-heavy cross-modal transformers for real-time surveillance applications.
Comments6 pages, 4 figures, 3 tables. Accepted at the 2026 IEEE International Conference on Advanced Video and Signal-based Surveillance (AVSS 2026). (c) 2026 IEEE. Personal use of this material is permitted. Permission from IEEE must be obtained for all other uses