EEG-Xplain:解码EEG基础模型的神经黑箱
EEG-Xplain: Decoding Neural Black-Boxes of EEG Foundation Models
浏览论文内容
中文总结 AI 辅助
针对EEG基础模型的黑箱问题,提出统一归因框架,整合多种解释方法分析空间、时间和频率维度,并结合LLM生成报告,实验验证解释与神经生理标记一致。
中文摘要 AI 辅助
EEG基础模型(如BIOT、LaBraM和EEGMamba)在神经信号解码中取得了显著性能,但其黑箱特性限制了临床信任和神经科学验证。我们提出了一个统一的归因框架,用于跨异构架构解释EEG基础模型。该框架整合了基于梯度、扰动和激活的解释方法,以在空间、时间和频率维度上分析模型行为。在空间上,它识别关键的EEG通道,并通过拓扑图可视化其分布。在时间上,它通过归因热图突出与决策相关的信号片段。在频率域中,它通过频谱扰动分析量化典型EEG节律的贡献。为评估解释可靠性,我们引入了一种结合扰动曲线下面积(AOPC)和跨方法一致性分析的人群级评估。该框架进一步利用大型语言模型(LLMs)将结构化归因输出转化为自然语言报告,弥合了低层神经表征与高层语义推理之间的鸿沟。在基准数据集(包括Mumtaz2016和TUAB)上的实验表明,生成的解释与既定的神经生理学标记一致,验证了有意义的神经表征,同时揭示了潜在的对伪影和虚假模式的依赖。所提出的框架为评估EEG基础模型的可解释性、可靠性和生理合理性提供了一种标准化方法。
英文摘要
EEG foundation models such as BIOT, LaBraM, and EEGMamba have achieved remarkable performance in neural signal decoding, but their black-box nature limits clinical trust and neuroscientific validation. We propose a unified attribution framework for interpreting EEG foundation models across heterogeneous architectures. The framework integrates gradient-, perturbation-, and activation-based explanation methods to analyze model behavior in spatial, temporal, and frequency dimensions. Spatially, it identifies critical EEG channels and visualizes their distributions using topographic maps. Temporally, it highlights decision-relevant signal segments through attribution heatmaps. In the frequency domain, it quantifies the contributions of canonical EEG rhythms via spectral perturbation analysis. To assess explanation reliability, we introduce a population-level evaluation combining Area Over the Perturbation Curve (AOPC) and cross-method consistency analysis. The framework further leverages Large Language Models (LLMs) to transform structured attribution outputs into natural-language reports, bridging low-level neural representations and high-level semantic reasoning. Experiments on benchmark datasets, including Mumtaz2016 and TUAB, demonstrate that the generated explanations are consistent with established neurophysiological markers, validating meaningful neural representations while exposing potential dependencies on artifacts and spurious patterns. The proposed framework provides a standardized approach for evaluating the interpretability, reliability, and physiological plausibility of EEG foundation models.
发表机构
- Guangzhou University(广州大学)
机构由 AI 辅助整理,请以论文原文为准。