arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

拓展早期诊断的视野:利用视觉Transformer进行肺癌预测

Extending the Horizon of Early Diagnosis: Lung Cancer Prediction with Vision Transformers

Olivera Kotevska, Ian Goethert, Michael McGee, Maria Mahbub, Sean R. Wilkinson, Rowena Yip, Myvizhi Esai Selvan, Zeynep H. Gumus, Claudia Henschke, Robert J. Klein, Providencia Morales, Samuel M Aguayo, Ioana Danciu, Mayanka Chandrashekar

arXiv 2608.21571首次发表:更新:

发表机构

Oak Ridge National Laboratory; Icahn School of Medicine at Mount Sinai; Phoenix VA Medical Center; Vanderbilt University Medical Center(橡树岭国家实验室; 西奈山伊坎医学院; 凤凰城退伍军人事务医疗中心; 范德堡大学医学中心)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究利用视觉Transformer(ViTs)分析259361张胸部X光片,通过解决类别不平衡优化模型,实现临床诊断前1-2年的肺癌预测,预训练ViTs性能优于从头训练模型,为早期肺癌风险预测提供了潜力。

AI 中文摘要

肺癌仍是全球癌症相关死亡的主要原因,早期诊断对提高生存率至关重要。然而,早期恶性肿瘤在胸部X光片上表现可能不明显,给放射科医生带来挑战。本研究评估视觉Transformer(ViTs)在临床诊断前1至2年预测肺癌的能力。我们分析了波士顿马萨诸塞州牙买加平原VA医院91020项影像研究中的259361张胸部X光片。该数据集存在极端类别不平衡,癌症与非癌症比例约为1:150,通过混合欠采样与过采样以及类别加权损失优化解决此问题。我们评估了三种ViT配置:从头训练的模型、在ImageNet上预训练的模型、在肺癌数据集上微调的Corona预训练模型。迁移学习提升了性能,预训练模型的AUC比从头训练的基线模型高出6-10个百分点,平衡准确率高出约10-12%。ImageNet预训练模型表现出最稳定的整体性能,而Corona预训练模型在某些场景下达到更高的灵敏度,但变异性更大。适度的重采样比例,包括1:1欠采样和1.5:2过采样,在灵敏度、精度和计算效率之间提供了良好的权衡,可将运行时间最多减少70%且无重大性能损失。这些发现证明了ViTs用于从常规胸部X光片进行早期肺癌风险预测的潜力。尽管性能仍低于临床部署阈值,但结果支持进一步开发基于ViT的分诊系统,以标记高风险患者进行更早评估。

英文摘要

Lung cancer remains a leading cause of cancer-related mortality worldwide, and early diagnosis is critical for improving survival. However, early-stage malignancies can be subtle on chest X-rays, creating challenges for radiologists. This study evaluates Vision Transformers (ViTs) for predicting lung cancer one to two years before clinical diagnosis. We analyzed 259,361 chest X-rays from 91,020 imaging studies at the Jamaica Plains VA Hospital in Boston, MA. The dataset showed extreme class imbalance, approximately 1:150 cancer to non-cancer, which was addressed using hybrid under- and over-sampling and class-weighted loss optimization. Three ViT configurations were evaluated: a model trained from scratch, an ImageNet-pretrained model, and a Corona-pretrained model fine-tuned on the lung cancer dataset. Transfer learning improved performance, with pretrained models exceeding the scratch baseline by 6-10 percentage points in AUC and about 10-12 percent in balanced accuracy. ImageNet-pretrained models showed the most stable overall performance, while Corona-pretrained models achieved higher sensitivity in some settings but greater variability. Moderate resampling ratios, including 1:1 undersampling and 1.5:2 oversampling, provided favorable trade-offs between sensitivity, precision, and computational efficiency, reducing runtime by up to 70 percent without major performance loss. These findings demonstrate the potential of ViTs for early lung cancer risk prediction from routine chest X-rays. Although performance remains below clinical deployment thresholds, the results support further development of ViT-based triage systems to flag high-risk patients for earlier evaluation.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑