arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

一种用于计算患者相似性的新型有监督训练方法

A New Trained Supervised Method for Calculating Patient Similarity

Minzee Kim, Joel A. Dubin

arXiv 2608.15973首次发表:更新:

AI 中文总结

本研究提出一种结合松弛自适应组套索的加权余弦相似性度量,用于二分类响应数据的个性化患者预测,经模拟及重症监护病房数据分析,其虽校准度略降但区分度大幅提升,整体预测性能更优。

AI 中文摘要

随着电子健康记录(Electronic Health Records)可用性的不断提高,个性化预测建模发展迅速。该方法旨在通过为每个个体拟合独特模型来提升模型的预测性能,我们基于与待预测个体具有相似性的训练数据子集训练模型,该相似性通过某种相似性度量识别。早期研究表明,使用基于定制数据子集训练的个性化模型,相比使用在完整数据集上训练的全局模型,能获得更好的预测效果。本研究针对二分类响应数据的个性化模型预测,提出一种新型患者相似性度量,具体引入加权余弦相似性度量,该度量在计算参与者间相似性时,通过分配预测变量特定权重扩展了标准余弦相似性,这些权重采用结合松弛自适应组套索(relaxed adaptive group lasso)的有监督方法估计。模拟研究及重症监护病房(intensive care unit)数据分析结果显示,尽管所提相似性度量会导致校准度略有下降,但在区分度上取得了显著提升;由于区分度的提升超过了校准度的损失,以Brier评分衡量的整体预测性能得到改善,因此,所提相似性度量能更有效地识别相似参与者,进而提升预测准确性。

英文摘要

Personalized predictive modelling has been growing rapidly with the increasing availability of Electronic Health Records. This approach aims to improve a model's predictive performance by fitting a unique model to each individual. We train the model on a subset of the training data consisting of individuals similar to the individual being predicted, identified through some similarity metric. Earlier studies show that using a personalized model trained on a customized subset of the data leads to better prediction than using a global model trained on the full dataset. In this work, we develop a new patient similarity metric to improve the prediction of a personalized model for binary response data. Specifically, we introduce a weighted cosine similarity metric that extends the standard cosine similarity metric by assigning predictor-specific weights when computing similarity between participants. These weights are estimated using a supervised approach with the relaxed adaptive group lasso. Results from simulation studies and an analysis of intensive care unit data show that although our proposed similarity metric leads to a slight deterioration in calibration, it produces substantial gains in discrimination. Overall predictive performance measured by the Brier Score improves because the increase in discrimination outweighs the loss in calibration; therefore, our proposed similarity metric more effectively identifies similar participants, resulting in improved predictive accuracy.

Comments26 pages

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑