arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

用于心力衰竭死亡率预测的离散时间生存分析

Discrete-Time Survival Analysis for Heart Failure Mortality Prediction

Aditya Rane, Amit Choudhari, Shashi Kant, Akash Deep

arXiv 2608.04140首次发表:更新:

AI 中文总结

本研究针对机器学习处理右删失生存数据的目标泄漏问题,提出离散时间人时框架,在UCI心力衰竭队列上验证了其有效性,发现随访时长不可作为基线预测因子,GAM模型表现最优。

AI 中文摘要

准确的心力衰竭预后依赖于随时间追踪临床风险,但许多机器学习应用处理右删失生存数据时存在问题,要么丢弃患者的观察时间,要么将其用作预测因子。丢弃时间会忽略生存背景,而将随访时间作为输入特征会引入严重的目标泄漏,从而夸大表观准确率。针对此问题,我们提出一种用于心力衰竭死亡率分类的离散时间人时框架。使用UCI心力衰竭临床记录队列(n=299,96例死亡),我们将数据转换为区间级二元结局,并将Cox比例风险基线模型与人时互补对数对数GLM和GAM模型,以及人时随机森林、XGBoost、随机生存森林和DeepSurv分类器进行基准测试。人时GLM重现了Cox风险比和一致性,验证了该转换的有效性;GAM捕捉到显著的非线性预测因子效应,在判别性和泛化性之间取得最佳平衡;灵活分类器虽表现出强劲的原始性能,但存在过拟合问题。最后,我们直接量化泄漏效应:纳入观察到的随访时长可将分类AUC从约0.73提升至近1.00,证实随访时长绝不能用作基线预测因子。总体而言,这些结果确立了一种生存感知框架,该框架将灵活分类与有效的事件发生时间结构相结合。

英文摘要

Accurate heart-failure prognosis relies on tracking clinical risk over time, yet many machine-learning applications mishandle right-censored survival data by either discarding a patient's observation time or using it as a predictor. Discarding time ignores survival context, while using follow-up time as an input feature introduces severe target leakage that inflates apparent accuracy. We address this by proposing a discrete-time person-period framework for heart-failure mortality classification. Using the UCI Heart Failure Clinical Records cohort ($n=299$, 96 deaths), we transform the data into interval-level binary outcomes and benchmark a Cox proportional hazards baseline against person-period complementary log-log GLM and GAM models, alongside person-period random forest, XGBoost, random survival forest, and DeepSurv classifiers. The person-period GLM reproduces the Cox hazard ratios and concordance, validating the transformation, while the GAM captures significant nonlinear predictor effects and provides the best balance of discrimination and generalization; the flexible classifiers achieve strong raw performance but overfit. Finally, we quantify the leakage effect directly, including observed follow-up duration raises classification AUC from roughly 0.73 to nearly 1.00, confirming that follow-up duration must not be used as a baseline predictor. Overall, these results establish a survival-aware framework that combines flexible classification with valid time-to-event structure.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑