AI 中文总结
针对参数化联合模型依赖强假设且易偏差的问题,提出无需参数假设的深度联合模型DeepAJM,利用编码器-解码器学习纵向轨迹并通过可解释关联结构提升生存预测,在三个数据集上取得了最佳判别性能。
AI 中文摘要
联合模型同时建模纵向和生存结局,利用患者纵向轨迹中的模式来改善生存结局的预测。然而,经典的参数化联合模型依赖于固定的参数假设,在模型设定错误和样本量较小的情况下容易产生偏差。我们提出了一种深度联合模型DeepAJM,它不需要任何参数假设,同时保留部分可解释的、按纵向结局划分的关联结构。该联合模型采用编码器-解码器(序列到序列)架构来学习患者时变协变量轨迹中的潜在结构。该模型通过一个可学习的可解释关联结构将纵向过程与生存过程联系起来,在该结构中,解码器的每个纵向输出在贡献于架构生存头的风险评分之前,会由基线协变量重新调制。该模型在三个数据集(一个心血管疾病电子健康记录队列、一个原发性胆汁性肝硬化(PBC2)数据集和一个模拟数据集)上,与经典参数化联合模型、TransformerJM、DA-LSTM以及一个基于Cox的仅生存模型进行了评估。所有模型均使用C指数、综合布里尔分数(IBS)、时间依赖性AUROC和时间依赖性AUPRC进行评估。我们的模型在所有数据集上就C指数、时间依赖性AUROC和AUPRC而言取得了最佳判别性能。
英文摘要
Joint Models simultaneously model longitudinal and survival outcomes, leveraging patterns in patients' longitudinal trajectory to improve the prediction of survival outcomes. The classical parametric joint models, however, rely on fixed parametric assumptions, making them susceptible to bias under model misspecification and smaller sample sizes. We propose a deep joint model, DeepAJM, that does not require any parametric assumptions, while retaining a partially interpretable, per-longitudinal-outcome association structure. The joint model uses an encoder-decoder (sequence-to-sequence) architecture to learn the latent structure in patients' time-varying covariate trajectories. The model links the longitudinal processes to the survival processes through a learned interpretable association structure, in which each longitudinal output from the decoder gets remodulated by baseline covariates before it contributes to the risk scores from the survival head of the architecture. The model was evaluated on three datasets ( a cardiovascular-disease EHR cohort, a primary biliary cirrhosis (PBC2) dataset, and a simulated dataset) against a classical parametric joint model, TransformerJM, DA-LSTM and a Cox-based survival-only model. All models were assessed using C-index, integrated brier score (IBS), time-dependent AUROC, and time-dependent AUPRC. Our model achieved the best discrimination in terms of the C-index, time-dependent AUROC, and AUPRC across all datasets.