发表机构
Leibniz Institute for Prevention Research and Epidemiology - BIPS; Institute of Science Tokyo; Faculty of Medicine Siriraj Hospital, Mahidol University(莱布尼茨预防研究与流行病学研究所——BIPS; 东京科学大学; 玛希隆大学诗里拉吉医学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
xMICD结合诊断分组与ICD嵌入相似性,构建可解释的患者表示,在多项临床预测任务中性能媲美ICD2Vec,兼具预测能力与临床可解释性。
AI 中文摘要
电子健康记录(EHRs)被广泛应用于机器学习驱动的临床风险预测,国际疾病分类(ICD)代码提供了患者诊断的结构化信息,但对其进行有效表示仍存在挑战。现有方法常面临预测性能与可解释性之间的权衡:基于分组的表示方法可解释性强,但可能丢失信息;而基于嵌入的表示方法能实现优异的预测性能,但难以解释。本文提出多ICD代码的可解释表示方法xMICD,用于从ICD代码集合构建低维患者表示。xMICD将具有临床意义的诊断分组与预训练ICD嵌入空间的相似性相结合,不使用二元分组隶属关系,而是通过基于相似性的相对分配将代码分配至各组,生成反映患者诊断与各临床组匹配紧密程度的特征。在大规模EHR数据集上的实验表明,xMICD在多项临床预测任务中实现了与ICD2Vec等基于嵌入的表示方法相当的预测性能,同时生成的特征仍具有临床可解释性,因为每个维度对应一个可识别的诊断组。因此,xMICD为将基于嵌入的语义关系整合到可解释的临床特征空间以用于机器学习模型提供了一种实用途径。
英文摘要
Electronic Health Records (EHRs) are widely used for clinical risk prediction using machine learning. International Classification of Diseases (ICD) codes provide structured information about patient diagnoses, but representing them effectively remains challenging. Existing approaches often face a trade-off between predictive performance and interpretability: grouping-based representations are interpretable but may lose information, while embedding-based representations achieve strong predictive performance but are difficult to interpret. We propose Explainable Representation of Multiple ICD Codes (xMICD), a method for constructing low-dimensional patient representations from sets of ICD codes. xMICD combines clinically meaningful diagnostic groupings with similarity in a pre-trained ICD embedding space. Instead of using binary group membership, the method assigns codes to groups via similarity-based relative assignments, yielding features that reflect how closely a patient's diagnoses align with each clinical group. Experiments on large-scale EHR datasets demonstrate that xMICD achieves predictive performance comparable to embedding-based representations such as ICD2Vec across multiple clinical prediction tasks. At the same time, the resulting features remain clinically interpretable because each dimension corresponds to a recognizable diagnostic group. xMICD therefore provides a practical way to integrate embedding-based semantic relationships into interpretable clinical feature spaces for machine learning models.