arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.07857cs.CVcs.AI

将CT基础模型蒸馏为可编辑概念瓶颈用于肺结节恶性程度预测

Distilling CT Foundation Models into Editable Concept Bottlenecks for Lung Nodule Malignancy Prediction

Fakrul Islam Tushar, Stephen Adamo, Geoffrey D. Rubin

首次发表
浏览论文内容

中文总结 AI 辅助

该研究将CT基础模型蒸馏为可编辑概念瓶颈模型,在肺结节恶性程度预测任务中,其判别力与仅用结节大小相当,且支持概念干预以实现透明可解释的预测。

中文摘要 AI 辅助

基础模型提供可迁移的CT表征,但基于这些嵌入的预测难以解释。我们开发了概念瓶颈模型,将两个冻结的CT基础模型表征映射到放射科医生定义的8个肺结节属性,并从估计的概念和结节大小预测恶性程度。这些模型包括CT-FM(一种使用96^3体素结节中心补丁的全CT自监督编码器)和FMCIB(一种使用50mm裁剪的结节聚焦对比编码器)。在2610个LIDC-IDRI结节上训练了8个岭回归概念头。恶性程度模型在LUNA25上训练,并在保留的内部测试集和外部DLCS队列上评估。概念保真度通过五折交叉验证的R^2评估,恶性程度判别力通过AUROC评估,95%置信区间通过按患者分组的自举重采样估计。概念保真度中等,但对于细微度(R^2,0.24 vs. 0.11)、毛刺征(0.17 vs. 0.08)、纹理(0.17 vs. 0.07)和分叶征(0.15 vs. 0.05),FMCIB高于CT-FM。内部测试中,CT-FM和FMCIB的概念+大小模型的AUROC分别为0.86(95% CI,0.80-0.92)和0.86(0.79-0.92);外部测试中,AUROC分别为0.72(0.68-0.75)和0.73(0.70-0.76),而仅结节大小的AUROC为0.73,对应仅嵌入探针的AUROC为0.60和0.67。加性预测可分解为特征级贡献,并通过受控概念干预修改。概念瓶颈提供透明的恶性程度预测,其判别力与仅结节大小相当,而概念保真度的差异表明概念恢复依赖于底层基础模型表征。

英文摘要

Foundation models provide transferable CT representations, but predictions based directly on these embeddings are difficult to interpret. We developed concept bottleneck models that map two frozen CT foundation-model representations to eight radiologist-defined pulmonary-nodule attributes and predict malignancy from the estimated concepts and nodule size. The models included CT-FM, a whole-CT self-supervised encoder using a 96^3-voxel nodule-centered patch, and FMCIB, a nodule-focused contrastive encoder using a 50-mm crop. Eight ridge-regression concept heads were trained on 2,610 LIDC-IDRI nodules. Malignancy models were trained on LUNA25 and evaluated on a held-out internal test set and the external DLCS cohort. Concept fidelity was assessed using five-fold cross-validated R^2, and malignancy discrimination was assessed using AUROC with 95% confidence intervals estimated by patient-grouped bootstrap resampling. Concept fidelity was modest but higher for FMCIB than CT-FM for subtlety (R2, 0.24 vs. 0.11), spiculation (0.17 vs. 0.08), texture (0.17 vs. 0.07), and lobulation (0.15 vs. 0.05). Internally, the CT-FM and FMCIB concept+size models achieved AUROCs of 0.86 (95% CI, 0.80-0.92) and 0.86 (0.79-0.92), respectively. Externally, AUROCs were 0.72 (0.68-0.75) and 0.73 (0.70-0.76), compared with 0.73 for nodule size alone and 0.60 and 0.67 for the corresponding embedding only probes. Additive predictions could be decomposed into feature-level contributions and modified through controlled concept interventions. Concept bottlenecks provided transparent malignancy predictions with discrimination similar to nodule size alone, while differences in concept fidelity suggest that concept recovery depends on the underlying foundation-model representation.

补充信息

↑