arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

个性化评分者建模:一种基于学习的框架,用于从多位专家处推导鲁棒的睡眠阶段标签

Personalized Scorer Modeling: A Learning-Based Framework for Deriving Robust Sleep Stage Labels from Multiple Experts

Seyyed Ali Hoseini, Javad Baseri, Hamid Saadatfar, Edris Hoseini Gol, AmirHossein Eshghi

arXiv 2608.12446首次发表:更新:

AI 中文总结

本研究提出基于学习的个性化评分者建模框架 LBH,利用多评分者数据集构建更可靠的睡眠阶段参考标签,在 DOD-H、DOD-O 数据集上较基线方法提升了睡眠分期性能。

AI 中文摘要

睡眠阶段分类对睡眠障碍的诊断和管理至关重要,但大多数自动分期研究在评估模型时仅采用单一参考 hypnogram( hypnogram 即 hypnogram,睡眠图),尽管已知评分者间存在变异性。本研究探究是否可利用多评分者数据集,从多位专家的集体行为中构建更可靠的参考标签。研究使用公开可用的 DOD-H 和 DOD-O 数据集,将 EEG(C3-M2)和颏肌电图(chin EMG)信号分割为 30 秒的 epoch(epoch 即 epoch,时间窗),并从每种模态中提取 30 个特征,EEG+EMG 模态共得到 60 个特征。我们提出一种基于学习的睡眠图(LBH),该方法利用机器学习模型得到的混淆矩阵,对每位评分者的阶段特定行为进行建模;经列归一化后,这些矩阵可估计每个真实睡眠阶段对应每位评分者标签的概率,再将多位评分者的概率聚合,为每个时间窗分配最终标签。LBH 在仅用 EEG 和 EEG+EMG 两种设置下,分别采用随机森林、支持向量机、多层感知机分类器进行评估,并与数据集睡眠图(DH)和最佳评分者睡眠图(BSH)对比。LBH 在所有情况下均提升了整体性能,在随机森林与 EEG+EMG 的组合下取得最佳结果:在 DOD-H 数据集上达到 86.07% 的准确率、85.46% 的精确率和 85.29% 的 F1 值,在 DOD-O 数据集上达到 86.04% 的准确率、85.21% 的精确率和 84.70% 的 F1 值。这些结果表明,个性化评分者建模可在不丢弃单个专家信息的前提下,提升参考睡眠图的构建质量。

英文摘要

Sleep stage classification is important for the diagnosis and management of sleep disorders, yet most automatic staging studies evaluate models against a single reference hypnogram despite known inter-scorer variability. This study investigates whether multi-scored datasets can be used to construct more reliable reference labels from the collective behavior of multiple experts. We use the publicly available DOD-H and DOD-O datasets. EEG (C3-M2) and chin EMG signals were segmented into 30-s epochs, and 30 features were extracted from each modality, yielding 60 features for EEG+EMG. We propose a learning-based hypnogram (LBH) that models the stage-specific behavior of each scorer using confusion matrices derived from machine-learning models. After column normalization, these matrices estimate the probability of each true sleep stage given each scorer's label; probabilities are aggregated across scorers to assign the final label for each epoch. LBH was evaluated with random forest, support vector machine, and multilayer perceptron classifiers under EEG-only and EEG+EMG settings, and compared with the dataset hypnogram (DH) and best-scorer hypnogram (BSH). LBH consistently improved overall performance. The best results were obtained with random forest and EEG+EMG, reaching 86.07% accuracy, 85.46% precision, and 85.29% F1-score on DOD-H, and 86.04% accuracy, 85.21% precision, and 84.70% F1-score on DOD-O. These findings suggest that personalized scorer modeling can improve reference hypnogram construction without discarding information from individual experts.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑