arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2507.07254eess.IVcs.CV

通过部分 CLIP 适配实现标签高效的胸部 X 射线诊断

Label-Efficient Chest X-ray Diagnosis via Partial CLIP Adaptation

Heet Nitinkumar Dalsania

更新

AI总结:

本文提出一种标签高效的胸部X射线诊断策略,通过对预训练CLIP视觉编码器进行部分微调,在NIH Chest X-ray14数据集上利用少样本学习有效提升了平均AUC得分,模拟了真实医院标注稀疏场景。

AI中文摘要:

医学影像的现代深度学习实现通常依赖大型标注数据集。由于隐私问题、高昂的成本甚至病例的稀缺性,这些数据集通常难以获取。本文提出了一种用于胸部 X 射线诊断的标签高效策略,旨在反映真实世界的医院场景。实验使用 NIH Chest X-ray14 数据集和预训练的 CLIP ViT-B/32 模型。通过对视觉编码器进行部分微调来适配模型,随后使用每种疾病类别 1-16 个标注样本进行零样本和少样本学习评估。测试表明,CLIP 的预训练视觉语言特征可有效适配少样本医学影像任务,与零样本基线相比,平均 AUC 得分提高了 20% 以上。本工作的关键在于尝试模拟医院内部工作流程,即存在图像档案但标注稀疏的情况。本工作评估了针对常见和罕见疾病诊断的实用且可扩展的解决方案。此外,本研究仅供学术和实验目的,尚未经过同行评审。所有代码可在 https://github.com/heet007-code/CLIP-disease-xray 找到。

英文摘要:

Modern deep learning implementations for medical imaging usually rely on large labeled datasets. These datasets are often difficult to obtain due to privacy concerns, high costs, and even scarcity of cases. In this paper, a label-efficient strategy is proposed for chest X-ray diagnosis that seeks to reflect real-world hospital scenarios. The experiments use the NIH Chest X-ray14 dataset and a pre-trained CLIP ViT-B/32 model. The model is adapted via partial fine-tuning of its visual encoder and then evaluated using zero-shot and few-shot learning with 1-16 labeled examples per disease class. The tests demonstrate that CLIP's pre-trained vision-language features can be effectively adapted to few-shot medical imaging tasks, achieving over 20\% improvement in mean AUC score as compared to the zero-shot baseline. The key aspect of this work is to attempt to simulate internal hospital workflows, where image archives exist but annotations are sparse. This work evaluates a practical and scalable solution for both common and rare disease diagnosis. Additionally this research is intended for academic and experimental purposes only and has not been peer reviewed yet. All code is found at https://github.com/heet007-code/CLIP-disease-xray.

↑