arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.20315cs.LG

用于结构化电子健康记录临床预测任务的可解释Transformer模型

Explainable Transformer Models for Clinical Prediction Tasks on Structured Electronic Health Records

Jun Ni Du, Lukas Adamek, Maxim Kryukov, Flavio Dormont, Ziv Bar-Joseph, Sven Jager, Brandon Rufino

首次发表
浏览论文内容

中文总结 AI 辅助

研究针对结构化EHR预测模型的可解释性缺口,提出BERT-LER模型,在EHRShot等任务中性能优于多数基准模型,归因符合临床风险因素,可推广至多领域。

中文摘要 AI 辅助

结构化电子健康记录(EHR)的预测模型仍是医疗机器学习的核心,但很少有模型同时强调定量实验室信息和关于输入医疗事件的可解释性。我们提出BERT-LER,一种基于BERT的编码EHR时间线模型,该模型从包含7500万患者的去标识化EHR数据集进行预训练和微调,将实验室检测结果编码为离散标记,同时通过基于百分位数的分箱保留分级信息,并结合集成梯度(Integrated Gradients)获取基于输入EHR序列的标记级归因。我们在公开的EHRShot基准套件和基于真实世界数据的哮喘严重程度进展研究中对该方法进行了基准测试,这通过在单一框架中统一实验室值表示和可解释性,解决了EHR基础模型建模中的方法学缺口,同时评估预测性能和解释是否能在标准临床预测任务之外实现泛化。在EHRShot和哮喘任务中,BERT-LER的预测性能与公开基准模型相当,且在实验室相关任务中通常优于这些模型,其提供的归因与临床已知的风险因素一致。我们的架构和可解释性方法可应用于使用结构化EHR训练的语言模型的多个治疗领域和预测任务。

英文摘要

Predictive models over structured electronic health records (EHRs) remain central to machine learning for healthcare, but few have jointly emphasized quantitative laboratory information and interpretability with respect to input medical events. We present BERT-LER, a BERT-style model for coded EHR timelines pretrained and fine-tuned from a de-identified EHR dataset of 75 million patients, that encodes laboratory test results as discrete tokens while retaining graded information through percentile-based binning, paired with Integrated Gradients for token-level attributions grounded in the input EHR sequence. We benchmark our approach on the public EHRShot benchmark suite and on an asthma severity progression study based on real-world data. This addresses a methodological gap in EHR foundation-style modeling by unifying laboratory value representation and explainability in a single framework, while assessing whether both predictive performance and explanations generalize beyond standard clinical prediction tasks. Across EHRShot and asthma tasks, BERT-LER achieves predictive performance that is competitive with, and on laboratory-related tasks often exceeds, publicly available benchmark models, and provides attributions that align with clinically known risk factors. Our architecture and explainability approach can be applied to many therapeutic areas and prediction tasks using language models trained on structured EHRs.

发表机构

  • Sanofi(赛诺菲)
  • Carnegie Mellon University(卡内基梅隆大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑