RAMamba-Net:一种用于听觉注意检测的可靠性感知与基于Mamba的多模态融合网络
RAMamba-Net: A Reliability-Aware and Mamba-Based Multimodal Fusion Network for Auditory Attention Detection
浏览论文内容
中文总结 AI 辅助
提出RAMamba-Net,一种可靠性感知的Mamba多模态融合网络,通过EEG-EOG融合和跨模态注意力提升听觉注意解码,在基准上获得5.76%准确率提升。
中文摘要 AI 辅助
听觉注意解码(AAD)从生理信号中识别被注意的说话者,支持神经引导的助听设备和自然的人机交互。脑电图(EEG)是AAD的主要模态,但在自然视听场景中提供的证据不完整,因此推动了脑电图和眼电图(EOG)的融合。现有方法仍受限于跨模态交互弱、时间建模效率低以及对样本变化鲁棒性差的问题。为解决这些局限,我们提出了RAMamba-Net,一种用于AAD的可靠性感知的基于Mamba的多模态融合网络。RAMamba-Net采用Mamba增强的频带感知卷积Transformer来捕获频带特定的EEG模式和长程时间动态。双分支时空编码器对EOG的时间依赖和通道间依赖进行建模。跨模态注意力实现了显式的模态交互。然后,引入可靠性感知模块来估计样本级别的模态权重,以实现特征和预测的一致性,从而增强多模态融合。在两个AAD基准上的实验表明,RAMamba-Net有效利用了EEG-EOG的互补信息,相比单模态基线获得了5.76%的准确率提升,同时具有更稳健的解码和更具判别性的表示。进一步的分析表明,显式的跨模态交互改善了多模态对齐,而可靠性感知模块抑制了不可靠的模态证据,并对信号扰动和参数变化具有鲁棒性。
英文摘要
Auditory attention decoding (AAD) identifies the attended speaker from physiological signals, supporting neuro-steered hearing devices and natural human-machine interaction. Electroencephalography (EEG) is the dominant modality for AAD but provides incomplete evidence in naturalistic audio-visual scenes, motivating EEG and electrooculography (EOG) fusion. Existing approaches remain limited by weak cross-modal interaction, inefficient temporal modeling, and low robustness to sample variations. To address the limitations, we propose RAMamba-Net, a reliability-aware Mamba-based multimodal fusion network for AAD. RAMamba-Net employs a Mamba-enhanced band-aware convolutional Transformer to capture band-specific EEG patterns and long-range temporal dynamics. A dual-branch temporal-spatial encoder models EOG temporal and inter-channel dependencies. Cross-modal attention enables explicit modality interaction. Then, a reliability-aware module is introduced to estimate sample-wise modality weights for feature and prediction consistency, thereby enhancing multimodal fusion. Experiments on two AAD benchmarks demonstrate that RAMamba-Net effectively exploits complementary EEG-EOG information, yielding accuracy gains of 5.76% over unimodal baselines, together with more robust decoding and discriminative representations. Further analyses show that explicit cross-modal interaction improves multimodal alignment, while the reliability-aware module suppresses unreliable modality evidence and is robust to signal perturbation and parameter variation.
发表机构
- School of Artificial Intelligence and Automation, Huazhong University of Science and Technology(华中科技大学人工智能与自动化学院)
机构由 AI 辅助整理,请以论文原文为准。