发表机构
University of Turku; Turku University Hospital; University College Cork; University at Albany, SUNY; Fudan University; University of Sydney(图尔库大学; 图尔库大学医院; 科克大学; 纽约州立大学奥尔巴尼分校; 复旦大学; 悉尼大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究比较六种睡眠分期模型在EEG、ECG及其组合下的性能,发现EEG承载主要信号,BIOT在EEG上最优(宏F1 0.7237),ECG替代平均损失0.3531,且预训练迁移有效。
AI 中文摘要
基于多导睡眠图(PSG)的自动睡眠分期是一个研究充分的任务,但PSG本身昂贵、基于临床且手动评分负担重,这限制了其在长期或居家监测中的使用。大多数现有的睡眠分期基础模型都是在完整的PSG导联配置下进行评估的。我们转而探究该导联配置中实际有多少是必要的。我们在动脉粥样硬化多种族研究(MESA)PSG数据集上评估了六种睡眠分期模型,涵盖三种信号条件:脑电图(EEG)、心电图(ECG)及其组合(EEG+ECG)。这一研究的动机源于边缘-云部署,其中EEG需要临床级头皮电极,而ECG已被消费级可穿戴设备捕获。我们测试了最先进的基础模型,如SleepFM(其编码器在MESA上从头训练),以及BIOT、MOMENT、LaBraM、SensorLM的基尺度视觉Transformer(ViT-B)重实现(在PyTorch中从头训练)和YASA,涵盖了EEG预训练、通用时间序列、从头训练和经典非学习方法。没有模型架构被修改其原始形式;SensorLM的编码器仅在PyTorch中重新实现。对于仅EEG分期,BIOT取得了最佳结果,宏F1为0.7237,其次是LaBraM(0.6835)和从头训练的SleepFM(0.6582)。在能够进行仅ECG分期的五个模型中,从EEG切换到ECG的宏F1代价介于0.2798(MOMENT)和0.4151(BIOT)之间,平均为0.3531,同时将原始通道数据速率削减至三分之一。对于大多数模型,将ECG添加到EEG并未带来增益。这些结果表明,EEG承载了大部分睡眠分期信号,量化了可穿戴兼容替代方案的一致准确性代价,并证明了睡眠相关预训练能很好地迁移到MESA。
英文摘要
Automatic sleep staging from polysomnography (PSG) is a well-studied task, but PSG itself is expensive, clinic-based, and burdensome to manually score, which limits its use for long-term or at-home monitoring. Most existing sleep-staging foundation models are evaluated using the full PSG montage. We instead ask how much of that montage is actually necessary. We evaluate six sleep staging models on the Multi-Ethnic Study of Atherosclerosis (MESA) PSG dataset across three signal conditions: electroencephalography (EEG), electrocardiography (ECG), and their combination (EEG+ECG). This is motivated by edge-cloud deployment, where EEG requires a clinic-grade scalp electrode, whereas ECG is already captured by consumer wearables. We test state-of-the-art foundation models such as SleepFM with an encoder trained from scratch on MESA, alongside BIOT, MOMENT, LaBraM, a base-scale Vision Transformer (ViT-B) reimplementation of SensorLM trained from scratch, and YASA, spanning EEG-pretrained, general-time-series, from-scratch, and classical non-learned approaches. No model architecture is modified from its original form; SensorLM's encoder is reimplemented only in PyTorch. For EEG-only staging, BIOT achieves the best result with a macro~F1 of 0.7237, followed by LaBraM (0.6835) and SleepFM from scratch (0.6582). Across the five models capable of ECG-only staging, switching from EEG to ECG costs between 0.2798 (MOMENT) and 0.4151 (BIOT) macro~F1, averaging 0.3531, while cutting the raw channel data rate to a third. Adding ECG to EEG provides no gain for most models. These results show that EEG carries most of the sleep-staging signal, quantify the consistent accuracy cost of the wearable-compatible alternative, and demonstrate that sleep-relevant pretraining transfers well to MESA.