arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

手术风险分层与结局预测中缺失的内容:端到端机器学习方法的范围综述

What Is Missing in Surgical Risk Stratification and Outcome Prediction: A Scoping Review of End-to-End Machine Learning Approaches

Yizhi Dong, Yuhe Ke, Hairil Rizal Abdullah, Yucheng Xing, Kevan Kai Bing Teo, Ling Huang, Mengling Feng

arXiv 2607.29090首次发表:更新:

AI 中文总结

本范围综述回顾190项研究,分析手术风险分层与结局预测的ML流程,发现数据、方法、评估等方面存在差距,为开发更严谨的围手术期ML工具提供参考。

AI 中文摘要

术后不良事件(包括死亡率和发病率)仍是全球主要负担,其中许多可通过早期识别高危患者和针对性围手术期护理预防,因此准确的风险分层至关重要。随着大规模电子健康记录(EHR)的可用性不断提高,机器学习(ML)提供了一种数据驱动的方法来建模复杂的临床模式。然而,现有研究的设计差异很大,方法实践仍不统一。本范围综述对使用EHR数据进行手术风险分层和结局预测的ML流程进行了特征分析。我们回顾了190项涵盖ML工作流程的研究,包括数据预处理、算法选择、模型评估和可解释性。大多数研究依赖于具有有限数据模态的单中心私有数据集,而开放获取手术数据集的稀缺限制了可重复性和可推广性。关键预处理步骤(包括缺失数据处理、特征选择和类别不平衡)的报告往往不完整。传统ML模型和简单神经网络占主导,而深度学习和多模态方法仍不常见。基准数据集和标准化评估协议基本缺失,阻碍了跨研究比较。仅约三分之一的研究纳入了可解释性方法。本综述确定了限制临床上稳健术后ML工具的方法学差距,并提供了结构化参考,以支持为围手术期护理开发更严谨、可重复且具有临床意义的ML。

英文摘要

Postoperative adverse events, including mortality and morbidity, remain a major global burden, many of which are preventable through early identification of high-risk patients and targeted perioperative care. Accurate risk stratification is therefore essential. With the growing availability of large-scale electronic health records (EHRs), machine learning (ML) provides a data-driven approach to model complex clinical patterns. However, existing studies vary widely in design, and methodological practices remain fragmented. This scoping review characterizes ML pipelines for surgical risk stratification and outcome prediction using EHR data. We reviewed 190 studies covering the ML workflow, including data preprocessing, algorithm selection, model evaluation, and explainability. Most studies relied on single-center private datasets with limited data modalities, while the scarcity of open-access surgical datasets constrained reproducibility and generalizability. Reporting of key preprocessing steps, including missing data handling, feature selection, and class imbalance, was often incomplete. Conventional ML models and simple neural networks predominated, whereas deep learning and multimodal approaches remained uncommon. Benchmark datasets and standardized evaluation protocols were largely absent, hindering cross-study comparisons. Only about one-third of studies incorporated explainability methods. This review identifies methodological gaps limiting clinically robust postoperative ML tools and provides a structured reference to support more rigorous, reproducible, and clinically meaningful ML development for perioperative care.

CommentsThis work has been submitted to the IEEE JBHI for possible publication. Copyright may be transferred without notice, after which this version may no longer be accessible

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑