arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.21384cs.AI

用于可泛化脑电信号表征学习的多模态预训练

Multimodal Pretraining for Generalizable EEG Representation Learning

Targol Bakhtiarvand, Jugal Kalita, Adham Atyabi

首次发表
浏览论文内容

中文总结 AI 辅助

研究针对癫痫EEG模型应用局限问题,开发多模态EEG基础模型,结合多种编码器,通过创新预训练技术创建癫痫相关表征。在多个数据集和评估设置下实验,该模型实现强大癫痫检测与适应新场景能力,支持可解释癫痫定位,性能达当前最优。

中文摘要 AI 辅助

用于癫痫的脑电图(EEG)模型通常局限于特定数据集和任务,难以跨数据集或不同场景应用。近期基础模型和自监督学习研究表明,适应性EEG主干可支持一系列EEG相关任务。本研究开发了一种多模态EEG基础模型,结合基于曼巴架构的原始信号编码器、用于时频数据的视觉Transformer(ViT)风格编码器和用于文本的轻量级编码器,均在共享嵌入空间内。预训练过程依赖多种创新技术,如掩码建模、跨视图对比对齐和时间一致性损失,无需标记数据即可创建丰富的癫痫相关表征。为评估预训练模型的有效性和泛化能力,在标准CHB-MIT癫痫检测基准和其他癫痫检测数据集上进行微调,并比较不同模型变体。在标准CHB-MIT分割上,最佳单模型AUROC达0.874,集成变体达0.878,代表该基准的当前最优性能。除标准训练-测试分割外,还在留一受试者(LOSO)协议下评估性能,19名受试者的平均LOSO平衡准确率为0.558。跨数据集和评估设置,多模态基础模型实现了强大的癫痫检测和对新癫痫检测场景的直接适应,同时支持可解释的癫痫定位。

英文摘要

Electroencephalography (EEG) models used for epilepsy are often limited to specific datasets and tasks. This limited approach can make it challenging to apply these models across different datasets or in various situations. However, recent studies in foundation models and self-supervised learning suggest that an adaptable EEG backbone could support a range of EEG related tasks. In this study, we have developed a multimodal EEG foundation model that combines a raw signal encoder based on the Mamba architecture, a Vision Transformer (ViT)-style encoder for time-frequency data, and a lightweight encoder for text, all within a shared embedding space. The pretraining process relies on several innovative techniques, such as masked modeling, cross-view contrastive alignment, and temporal consistency losses. These methods are designed to create rich, seizure-relevant representations without requiring labeled data. To assess the efficacy and generalization of our pretrained model, we fine-tuned it on the canonical CHB-MIT seizure detection benchmark and additional seizure detection datasets, and conducted extensive experiments comparing different model variants. On the standard CHB-MIT split, our best single model achieved an AUROC of 0.874, and an ensemble variant reached 0.878 AUROC, representing state-of-the-art performance on this benchmark. In addition to standard train-test splits, we evaluated performance under a leave-one-subject-out (LOSO) protocol, which is rarely reported in prior EEG seizure modeling work and highlights the difficulty of patient-independent seizure detection, with a mean LOSO balanced accuracy of 0.558 across 19 subjects. Across datasets and evaluation settings, our multimodal foundation model enabled robust seizure detection and straightforward adaptation to new seizure detection scenarios, while also supporting interpretable seizure localization.

发表机构

  • University of Colorado Colorado Springs(科罗拉多大学科罗拉多斯普林斯分校)

机构由 AI 辅助整理,请以论文原文为准。

↑