发表机构
Microsoft Research; Providence Genomics; Earle A. Chiles Research Institute, Providence Cancer Institute; The Oregon Clinic(微软研究院; 普罗维登斯基因组学; 厄尔·A·奇尔斯研究所,普罗维登斯癌症研究所; 俄勒冈诊所)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究提出Survprompt框架,将患者协变量转为临床文本提示,使零样本LLM在生存预测上接近专门模型,但跨癌症类型和机构性能不稳定,限制了临床使用。
AI 中文摘要
生存分析根据患者协变量估计事件发生时间的结果,并广泛用于医学风险评估。患者在诊断后寻求预后信息时,可能会求助于大型语言模型(LLM),这些模型现在通过消费级应用易于获取。然而,LLM能否提供准确的生存预测尚未经过严格评估。我们引入了Survprompt框架,该框架将结构化患者协变量转换为自由文本临床小插曲,并提示预训练的LLM进行零样本生存预测。我们将Survprompt与包括随机生存森林(RSF)在内的传统生存模型进行了基准测试,跨越两个多机构泛癌队列:公开可用的MSK-CHORD队列和新整理的普罗维登斯圣约瑟夫健康网络队列,后者使用基于LLM的医学抽象框架构建。我们报告了删失平均绝对误差(cMAE)和一致性指数(c-index),并进行特征消融以识别影响LLM预测的变量。前沿LLM在个体生存时间方面取得了令人惊讶的竞争性cMAE。例如,GPT-5.6-Sol在MSK-CHORD中针对几种癌症类型实现了cMAE在专门为生存预测训练的最先进RSF模型的10%以内,并且对于前列腺癌的cMAE低于RSF。特征消融显示,LLM优先考虑的临床变量与专门生存模型相似。然而,LLM在不同癌症类型和机构间表现出不一致的准确性,并且在高风险与低风险患者的区分上表现不佳(较低的c-index)。零样本LLM无需专门训练即可生成令人惊讶的准确预后估计,但其在不同癌症类型和机构间的可变性能仍然是临床使用的重要限制。
英文摘要
Survival analysis estimates time-to-event outcomes from patient covariates and is widely used for medical risk assessment. Patients seeking prognostic information after a diagnosis may turn to large language models (LLMs), now readily accessible through consumer applications. However, whether LLMs can provide accurate survival predictions has not been rigorously evaluated. We introduce Survprompt, a framework that converts structured patient covariates into free-text clinical vignettes and prompts pre-trained LLMs to predict survival zero-shot. We benchmark Survprompt against conventional survival models, including random survival forests (RSF), across two multi-institutional pan-cancer cohorts: the publicly available MSK-CHORD cohort and a newly curated cohort from the Providence St. Joseph Health Network constructed using an LLM-based medical abstraction framework. We report censored mean absolute error (cMAE) and concordance index (c-index) and conduct feature ablations to identify variables influencing LLM predictions. Frontier LLMs achieved surprisingly competitive cMAE for individual survival times. For example, GPT-5.6-Sol achieved cMAE within 10% of state-of-the-art RSF models specifically trained for survival prediction for several cancer types and lower cMAE than RSF for prostate cancer in MSK-CHORD. Feature ablations revealed that LLMs prioritized clinical variables similarly to specialized survival models. However, LLMs showed inconsistent accuracy across cancer types and institutions and poorly discriminated between high- and low-risk patients (lower c-index). Zero-shot LLMs can generate surprisingly accurate prognostic estimates without specialized training, but their variable performance across cancer types and institutions remains an important limitation for clinical use.