Foresight-England:面向COVID-19大流行期间医疗事件预测的全国规模电子健康记录生成式AI模型的开发
Foresight-England: Development of a National-Scale Generative AI Model of Electronic Health Records for Medical Event Prediction across the COVID-19 Pandemic
浏览论文内容
中文总结 AI 辅助
本研究开发了首个全国规模的电子健康记录生成式基础模型Foresight-E,在6100万患者数据上训练,可零样本预测医疗事件,为大流行期间的医疗事件预测提供了方法模板。
中文摘要 AI 辅助
Foresight-England(Foresight-E)是首个严格针对COVID-19研究开发的全国规模电子健康记录(EHR)生成式基础模型,我们评估了其对大流行直接和间接影响的建模能力。Foresight-E完全在英格兰国民保健署(NHS)安全数据环境中从头开始训练,是一个拥有2.43亿参数的Transformer解码器模型。它在约6100万名个体的去标识化纵向EHR上进行训练和评估,整合了初级/二级医疗保健数据、死亡登记数据和COVID-19数据;训练和验证使用了2018年11月至2022年12月期间的90%子集(5490万人),剩余10%(610万人)被留出用于评估。Foresight-E以自回归方式建模患者时间线,根据患者既往病史预测下一个医疗事件;在推理阶段,它采用零样本方式运行,无需针对特定任务进行训练即可预测其约40000个编码词汇中的任意概念。我们的分词方案保留了ICD-10、OPCS-4和SNOMED CT编码的临床粒度,共同表示绝对和相对时间。我们设计了一个针对COVID-19相关的30天住院和死亡的评估框架,包括按人口统计学因素和疫苗接种状态进行的亚组分析;为评估其对未见过的未来数据的泛化能力以及大流行的间接影响,我们在2023年(超出其训练期)的医疗事件上对该模型进行了测试,并与逻辑回归和XGBoost进行了基准对比。根据项目状态部分的详细信息,英格兰NHS已暂停Foresight-E项目的数据访问,因此目前无法获得定量结果;相反,我们分享了分词、架构、训练、推理和评估的策略,作为构建人群规模EHR基础模型所面临挑战的方法模板和案例研究。
英文摘要
Foresight-England (Foresight-E) is the first national-scale generative foundation model of electronic health records (EHRs), developed as a research pilot strictly for COVID-19 research. We evaluated its ability to model the direct and indirect effects of the pandemic. Trained from scratch entirely within the NHS England Secure Data Environment, Foresight-E is a 243-million-parameter transformer decoder. It was trained and evaluated on de-identified, longitudinal EHRs of approximately 61 million individuals, integrating primary/secondary care, death registrations, and COVID-19 data. Training and validation used a 90% subset (54.9 million) spanning November 2018 to December 2022; the remaining 10% (6.1 million) was held out for evaluation. Foresight-E models patient timelines autoregressively, predicting the next medical event given their prior history. At inference, it operates zero-shot, predicting any concept in its ~40,000-code vocabulary without task-specific training. Our tokenisation scheme retains the clinical granularity of ICD-10, OPCS-4, and SNOMED CT codes, jointly representing absolute and relative timing. We designed an evaluation framework for 30-day COVID-19 hospitalisation and mortality, including subgroup analyses by demographic factors and vaccination status. To assess generalisation to unseen future data and the pandemic's indirect effects, we tested the model on medical events from 2023 (beyond its training period), benchmarking against logistic regression and XGBoost. As detailed in the Project Status section, NHS England has paused access to data for the Foresight-E project, meaning quantitative results are currently unavailable. Instead, we share our strategy for tokenisation, architecture, training, inference, and evaluation as a methodological template and case study in the challenges of building population-scale EHR foundation models.
发表机构
- University College London(伦敦大学学院)
- King’s College London(伦敦国王学院)
- University College London Hospitals National Institute for Health Research Biomedical Research Centre(伦敦大学学院医院国家卫生研究院生物医学研究中心)
- Interdisciplinary Transformation University(跨学科转型大学)
- British Heart Foundation Data Science Centre(英国心脏基金会数据科学中心)
- Health Data Research UK(英国健康数据研究中心)
- The University of Edinburgh(爱丁堡大学)
- University of Cambridge(剑桥大学)
- Victor Phillip Dahdaleh Heart and Lung Research Institute, University of Cambridge(剑桥大学维克多·菲利普·达德赫心肺研究所)
- British Heart Foundation Centre of Research Excellence, University of Cambridge(剑桥大学英国心脏基金会卓越研究中心)
机构由 AI 辅助整理,请以论文原文为准。