arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

修复一个学习到癌症更严重意味着风险更低的模型:膀胱癌复发预测中的单调约束

Fixing a Model That Learned Worse Cancer Means Lower Risk: Monotonic Constraints in Bladder Cancer Recurrence Prediction

Saram Abbas, David Thomas, Naeem Soomro, Rishad Shafik, Rakesh Heer, Kabita Adhikari

arXiv 2610.00858首次发表:更新:

发表机构

Newcastle University; University of Southampton; Freeman Hospital; Imperial College London(纽卡斯尔大学; 南安普顿大学; 弗里曼医院; 帝国理工学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究针对膀胱癌复发预测中模型学习到与医学直觉相反的关系(更严重癌症预测更低风险)的问题,提出反事实方向测试和单调约束框架,在不损失性能的情况下检测并修复此缺陷,确保模型符合临床预期。

AI 中文摘要

背景与目标:临床医生预期复发风险随癌症严重程度升高而增加。在一项英国多中心试验中,一个无约束的XGBoost模型学习到更高的肿瘤分期和原位癌预测更低的复发风险,而判别力、校准和SHAP均未发现此问题。我们开发了一个反事实测试框架来检测这种反转,以及一个单调约束框架来在不损害性能的情况下消除它。方法:BOXIT试验在51个英国中心(2007-2012年)招募了472名接受方案规定膀胱镜检查的患者;其中435名患者至少有两年随访(153例复发,占35.2%)。我们开发了一个反事实方向测试和单调约束修正,约束方向来自EORTC和EAU风险系统,并在50折交叉验证中,使用18个预测因子(其中7个有方向约束),将两者与无约束XGBoost和逻辑回归进行了评估。该测试每次在一个有方向约束的特征上恶化每个患者,以检查风险是否下降;同时评估了SHAP方向和校准。主要发现与局限性:肿瘤分期和原位癌与更低的复发风险相关,与医学直觉相反;无约束模型在90.2%的病例中反转了原位癌的反事实,在74.3%的病例中反转了分期的反事实。判别力(ΔAUC 0.005,p=0.47)、校准和SHAP幅度均未发现此反转。单调约束在不对判别力造成损失(0.723对0.718)的情况下消除了所有违规,并优于EORTC(p=8.9e-16)。局限性:单一试验,仅进行了内部-外部验证。结论与临床意义:一个学习了这种反转的模型通过了所有常规检查。反事实方向测试,作为一次带有预指定单调约束的重新拟合运行,能在不损失性能的情况下捕获此失败,应在临床部署前常规进行。

英文摘要

Background and Objective: Clinicians expect recurrence risk to climb with cancer severity. In a UK multicentre trial, an unconstrained XGBoost model learnt that higher tumour stage and carcinoma in situ predicted lower recurrence risk, and discrimination, calibration, and SHAP were all blind to it. We developed a counterfactual testing framework to detect this inversion and a monotonic-constraint framework to remove it without hurting performance. Methods: BOXIT enrolled 472 patients with protocol-mandated cystoscopy across 51 UK sites (2007-2012); 435 had at least two years' follow-up (153 recurrences, 35.2%). We developed a counterfactual direction test and a monotonic-constraint correction, with constraint directions drawn from the EORTC and EAU risk systems, and evaluated both against unconstrained XGBoost and logistic regression on 18 predictors (seven directed) over 50 cross-validation folds. The test worsened each patient on one directed feature at a time to check whether risk fell; SHAP direction and calibration were also assessed. Key Findings and Limitations: Tumour stage and carcinoma in situ were associated with lower recurrence, opposite to medical intuition; the unconstrained model reversed carcinoma in situ counterfactuals in 90.2% of cases and stage in 74.3%. Discrimination ($Δ$AUC 0.005, p=0.47), calibration, and SHAP magnitude were all blind to the inversion. Monotonic constraints eliminated every violation at no cost to discrimination (0.723 vs 0.718) and outperformed EORTC (p=8.9e-16). Limitations: single trial, internal-external validation only. Conclusions and Clinical Implications: A model that had learned this inversion passed every conventional check. A counterfactual direction test, run as a single refit with pre-specified monotonic constraints, catches this failure at no cost to performance and should be routine before clinical deployment.

Comments13 pages, 4 figures, 1 table. Supplementary material (15 pages) provided as ancillary files

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑