发表机构
Université Laval; NeuroGenis Inc.(拉瓦尔大学; NeuroGenis公司)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究预训练脑电图基础模型在临床解码中的表现,通过多种方法在多数据集多任务上对六个模型进行基准测试,发现模型结论依赖评估单元、数据集偏移等因素,如在韩国痴呆症等任务中不同模型和方法表现各异。
AI 中文摘要
预训练的脑电图基础模型越来越多地被用于临床解码,但其在不同人群中的迁移能力以及对阴性对照的稳健性仍不明确。我们使用冻结线性探针,通过留一受试者法、按受试者分组或明确识别记录级分割,在四个数据集上的五个临床任务中对六个模型(LaBraM、EEGMamba、CBraMod、REVE、BENDR和BIOT)进行基准测试,并对选定的REVE结果进行随机初始化、随机特征、标签排列、加扰标签微调及投影敏感性测试。在韩国痴呆症(CAUEEG,三分类)任务中,冻结的REVE模型的AUROC为0.568,而经典特征为0.769;在患者不相交的保留分割中排序依然如此(0.565对0.768)。数据集标识可从冻结嵌入中轻松解码(PCA-50时AUROC为1.000;波段限制和逐epoch z评分后为0.9998),而相同的PCA-50管道对韩国诊断的解码AUROC为0.528。在该任务上,随机初始化的编码器也优于预训练的REVE(0.659对0.570)。在阿尔茨海默病任务中,相同预训练嵌入的高斯随机投影和PCA表现相似,经典特征在受试者层面名义上超过REVE。最明显的受控阳性是在CHB-MIT(n=23)上的跨受试者发作期检测,REVE的AUROC为0.793,比随机初始化的编码器高9.2个百分点。这些结果表明,脑电图基础模型的结论强烈依赖于评估单元、数据集偏移、比较器强度和目标对照。
英文摘要
EEG foundation-model gains may depend on cohort, montage, or probe design. We evaluated five models on five tasks across four benchmark datasets plus Korean CAUEEG, using subject-disjoint validation where identifiers exist. CAUEEG is recording-level with an annotated no-overlap held-out sensitivity. On matched CAUEEG normal/mild cognitive impairment/dementia classification (1,187 recordings), classical features reached 0.734 macro-AUROC (enhanced sensitivity: 0.736), versus BIOT-bipolar16 0.677, CBraMod 0.669, and REVE 0.568. The annotated no-overlap held-out subset preserved the classical-over-REVE ordering (0.717 versus 0.565). All five encoders decoded dataset identity at 1.000 before and after in-fold PCA-50; label permutations collapsed to chance and balanced subsamples remained at 1.000. This establishes dataset membership, not a causal site, geography, or population effect. A matched fully randomly initialized encoder was descriptively higher than pretrained REVE on CAUEEG (0.667 versus 0.570), and correct- versus scrambled-source-label LoRA runs yielded numerically similar AUROCs in unmatched descriptive sensitivities, not label-effect estimates or equivalence tests. On CHB-MIT cross-subject ictal detection, REVE reached 0.793 AUROC, versus 0.739 for the best tested enhanced nonlinear comparator, 0.691 for fully random initialization, and 0.505 for raw-signal random features. The paired REVE-minus-enhanced-comparator difference was +5.38 percentage points (95% CI -0.36 to +11.22), so comparator superiority remains unresolved; an amplitude-aware comparator also cannot be reconstructed from the retained normalized inputs. We distill montage matching, patient-overlap checks, stronger comparators, and representation controls into a reporting protocol for clinical EEG foundation-model studies.