CACSurv:基于大型语言模型的一致性对齐比较学习用于癌症生存预测
CACSurv: Concordance-Aligned Comparative Learning with Large Language Models for Cancer Survival Prediction
浏览论文内容
中文总结 AI 辅助
该研究针对癌症生存预测中LLM时间回归的公式与监督不匹配问题,提出CACSurv框架,构建TCGA-SurvReport基准,在6个TCGA癌症队列上的平均C指数达0.722,显著优于现有模型与基线。
中文摘要 AI 辅助
癌症生存预测支持治疗规划、风险分层和随访管理。现有方法使用结构化临床变量、全切片图像、基因组谱或多模态输入,而患者报告仍未得到充分探索。我们研究以报告为中心的生存预测,使用的报告整合了病理、临床和分子证据。大型语言模型(LLMs)可对这类报告进行推理,但逐例时间回归会引发两种不匹配:一是公式不匹配,因为生存评估依赖于可比患者的排序,而独立时间预测不强制排序一致性;二是监督不匹配,因为删失患者的观察时间表明其存活超过该时间点,无法作为精确的回归目标,但仍暗示了其与更早死亡患者的排序关系。为解决这些不匹配,我们提出CACSurv,这是一个用于以报告为中心的生存预测的一致性对齐比较框架。CACSurv将生存建模重新表述为小型队列比较推理,其中LLM预测相对预后排序。我们引入基于右删失下可比关系的一致性对齐奖励,使删失结果能提供排序监督,无需精确事件时间目标。推理时,蒙特卡洛参考聚合将每个患者与采样参考进行比较,并将位置聚合成队列级排序。我们建立了TCGA-SurvReport,这是一个涵盖6个TCGA癌症队列的基准。CACSurv在所有6个队列上均达到最高C指数,平均C指数为0.722,比最强的已发表生存模型高出6.5个百分点,比最强的LLM时间回归基线高出4.2个百分点。我们的代码、模型和数据集将在this https URL处公开。
英文摘要
Cancer survival prediction supports treatment planning, risk stratification, and follow-up management. Existing methods use structured clinical variables, whole-slide images, genomic profiles, or multimodal inputs, while patient reports remain underexplored. We study report-centric survival prediction using reports that organize pathological, clinical, and molecular evidence. Large language models (LLMs) can reason over such reports, but case-wise time regression introduces two mismatches. First, a formulation mismatch arises because survival evaluation depends on ordering comparable patients, whereas independent time predictions do not enforce ranking consistency. Second, a supervision mismatch arises because a censored patient's observed time indicates survival beyond that point and cannot serve as an exact regression target, although it still implies orderings relative to patients who died earlier. To address these mismatches, we propose CACSurv, a Concordance-Aligned Comparative framework for report-centric survival prediction. CACSurv reformulates survival modeling as mini-cohort comparative reasoning, where an LLM predicts relative prognostic orderings. We introduce concordance-aligned rewards derived from comparable relations under right censoring, enabling censored outcomes to provide ranking supervision without exact event-time targets. At inference, Monte Carlo Reference Aggregation compares each patient with sampled references and aggregates positions into a cohort-level ranking. We establish TCGA-SurvReport, a benchmark covering six TCGA cancer cohorts. CACSurv achieves the highest C-index on all six cohorts and an average C-index of 0.722, outperforming the strongest published survival model by 6.5 percentage points and the strongest LLM time-regression baseline by 4.2 percentage points. Our code, models, and dataset will be available at https://github.com/xmed-lab/CACSurv.