arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.09431stat.COstat.ML

Comprisk:用于竞争风险生存分析的与scikit-learn兼容的Python工具包

comprisk: A scikit-learn-compatible Python toolkit for competing-risks survival analysis

Sunny Yang, Weiyan Zhao, Wanqi Zhao

首次发表
浏览论文内容

中文总结 AI 辅助

针对医疗事件时间数据竞争风险问题,Comprisk工具包整合多种竞争风险方法并提供一致API,增加模型评估,各估计器经数值验证,森林模型速度快且可扩展,能让研究人员在Python科学栈内进行竞争风险分析。

中文摘要 AI 辅助

医疗事件发生时间数据常面临竞争风险,标准生存方法会产生有偏差的绝对风险估计。正确分析应针对特定原因累积发病率函数(CIF)。以往应用研究人员多通过R包进行此分析,迫使基于Python的机器学习工作流程在Python和R之间往返。本文提出Comprisk工具包,它整合多种竞争风险方法,拥有一致API,还增加了相关模型评估。每个估计器都经过数值验证,森林模型速度更快且可扩展,该工具包可让研究人员在Python科学栈内进行正确且可扩展的竞争风险分析。

英文摘要

Medical time-to-event data are frequently subject to competing risks, where the occurrence of one terminal event precludes the others and standard survival methods that treat competing events as censoring yield biased absolute-risk estimates. Valid analysis instead targets the cause-specific cumulative incidence function (CIF). This methodology has been available to applied researchers almost exclusively through R packages, forcing Python-based machine-learning workflows into a Python-to-R round trip. We present comprisk, a scikit-learn-compatible Python toolkit that puts the canonical competing-risks methods behind one API: a scalable competing-risks random survival forest, Fine-Gray subdistribution-hazard regression and a penalized variant, cause-specific Cox regression, the Aalen-Johansen CIF estimator, and Gray's K-sample test, together with competing-risks-aware model evaluation. Every estimator is validated numerically against its R reference implementation. The forest uses a histogram-based, numba-compiled split kernel that fits 10-22x faster than randomForestSRC at comparable discrimination on real clinical cohorts and scales to n = 10^6 on a consumer CPU. comprisk is distributed on PyPI and lets applied researchers run correct, scalable competing-risks analysis without leaving the Python scientific stack.

补充信息

↑