发表机构
Hopewell Valley Central High School(霍普韦尔谷中央高中)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究审计EEG基础模型在运动想象任务上的表现,发现匹配输入下不同架构的精度差异符号相反,表明单一比较器无法提供架构不变的性能差距分解。
AI 中文摘要
预训练的脑电(EEG)基础模型越来越多地被提出作为脑机接口的通用编码器,然而最近的基准测试对其表示何时能迁移到下游任务存在分歧。我们在一个验证锁定的协议下审计了LaBraM和CBraMod在运动想象任务上的表现,该协议中预处理、架构、优化、冻结深度、检查点、温度和方法选择仅使用训练会话数据来确定。在四类BCI竞赛IV-2a上,这里评估的每个有监督比较器都优于每个基础模型配置,包括验证选择的微调。然后我们检查了一个关键的混淆因素:基础模型和任务特定解码器通常使用不同的输入流程进行评估。在基础模型消耗的宽带阵列上重新训练三个有监督架构,产生了跨架构符号相反的匹配输入精度差异:宽带输入使ATCNet的精度提高了0.078,同时使EEG Conformer的精度降低了0.088。在n=9时经过多重比较校正后,三个单独的匹配输入项均不显著,因此我们将符号变化视为描述性的,而非正式的架构-流程交互。这些观察到的符号差异表明,单个比较器可能无法提供预训练与有监督性能差距的架构不变分解。四类缺陷也不会在运动想象数据集上均匀重现:在两类BNCI2014-004上,我们无法检测到微调后的CBraMod与有监督比较器之间的相同分离。最后,验证拟合的温度缩放将基础模型的校准误差恢复到有监督范围内,尽管四类精度显著较低。
英文摘要
Pretrained EEG foundation models are increasingly proposed as general-purpose encoders for brain-computer interfaces, yet recent benchmarks disagree about when their representations transfer to downstream tasks. We audit LaBraM and CBraMod on motor imagery under a validation-locked protocol in which preprocessing, architecture, optimization, freeze depth, checkpoint, temperature, and method selection are determined using training-session data only. On four-class BCI Competition IV-2a, every supervised comparator evaluated here outperforms every foundation-model configuration, including validation-selected fine-tuning. We then examine a key confound: foundation models and task-specific decoders are normally evaluated using different input pipelines. Retraining three supervised architectures on the broadband arrays consumed by the foundation models produces matched-input accuracy differences of opposite sign across architectures: broadband input improves ATCNet by 0.078 accuracy while reducing EEG Conformer accuracy by 0.088. None of the three individual matched-input terms is significant after multiple-comparison correction at n = 9, so we treat the sign variation descriptively rather than as a formal architecture-by-pipeline interaction. These observed sign differences suggest that a single comparator may not provide an architecture-invariant decomposition of a pretrained-versus-supervised performance gap. The four-class deficit also does not reproduce uniformly across motor-imagery datasets: on two-class BNCI2014-004 we cannot detect the same separation between fine-tuned CBraMod and the supervised comparators. Finally, validation-fitted temperature scaling returns foundation-model calibration error to the supervised range despite substantially lower four-class accuracy.