arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

Med-AR:用于长尾胸部X射线分类和不确定性感知评估的自回归视觉-语言预训练

Med-AR: Autoregressive Vision-Language Pretraining for Long-Tailed Chest X-Ray Classification and Uncertainty-Aware Evaluation

Janhavi Prabhu, Sahil, Akshay V, Shivam Shukla, Manoj Tadepalli, Preetham Putha

arXiv 2609.29156首次发表:更新:

AI 中文总结

本文提出Med-AR自回归视觉-语言预训练模型,在长尾胸部X射线分类中优于对比学习方法,提升尾部标签AUPRC,并改善选择性预测性能。

AI 中文摘要

长尾胸部X射线分类需要能够同时捕捉常见异常和细微、罕见发现的视觉表征。我们提出了Med-AR-8B和Med-AR-2B,两个放射学原生的自回归视觉-语言模型,使用结构化报告、异常聚焦文本和区域标注进行预训练。我们评估了其视觉编码器在多标签分类中的迁移性能,与对比学习、自监督和监督预训练编码器(包括Med-CLIP、CheXFound、EVA-Base、ARK和BioViL-T)进行比较,使用统一的ML-Decoder分类头。为了评估细粒度识别能力,我们还为MIMIC-CXR和CheXpert构建了由LLM扩展的、基于报告的标签集。在PadChest、MIMIC-CXR和CheXpert上,Med-AR-8B在头部、中部和尾部发现的平均AUROC和AUPRC上均优于Med-CLIP。在MIMIC-CXR上,它将尾部标签的平均AUPRC从0.1033提高到0.1441。Med-AR-2B在PadChest上取得了最强的判别结果。在更广泛的编码器比较中,Med-AR的一个变体在每个公开数据集的每个报告患病率组中均取得了最高的平均AUROC和AUPRC。两个Med-AR变体在所有三个公开数据集上的风险覆盖曲线下超额面积也低于Med-CLIP,表明在所评估的协议下具有改进的选择性预测性能。内部结果依赖于指标,Med-CLIP在总体和尾部AUPRC以及选择性预测方面保持优势。这些发现确立了Med-AR作为在评估的公开基准上进行长尾胸部X射线分类的强预训练方案,并展示了同时评估判别能力和选择性预测的价值。

英文摘要

Long-tailed chest X-ray classification requires visual representations that capture both common abnormalities and subtle, infrequent findings. We propose Med-AR-8B and Med-AR-2B, two radiology-native autoregressive vision-language models pretrained with structured reports, abnormality-focused text, and region annotations. We evaluate the transfer of their visual encoders to multi-label classification against contrastive, self-supervised, and supervised pretrained encoders, including Med-CLIP, CheXFound, EVA-Base, ARK, and BioViL-T, using a common ML-Decoder classification head. To assess fine-grained recognition, we also construct LLM-expanded, report-derived label sets for MIMIC-CXR and CheXpert. Across PadChest, MIMIC-CXR, and CheXpert, Med-AR-8B outperforms Med-CLIP in mean AUROC and AUPRC for head, medium, and tail findings. On MIMIC-CXR, it increases tail-label mean AUPRC from 0.1033 to 0.1441. Med-AR-2B achieves the strongest discrimination results on PadChest. Across the broader encoder comparison, a Med-AR variant achieves the highest mean AUROC and AUPRC in every reported prevalence group on each public dataset. Both Med-AR variants also achieve lower excess area under the risk-coverage curve than Med-CLIP on all three public datasets, indicating improved selective-prediction performance under the evaluated protocol. Internal results are metric-dependent, with Med-CLIP retaining advantages in overall and tail AUPRC and in selective prediction. These findings establish Med-AR as a strong pretraining recipe for long-tailed chest X-ray classification on the evaluated public benchmarks and demonstrate the value of assessing discrimination and selective prediction together.

Comments80 pages including supplementary material, 28 figures, and 22 tables. Supplementary material is included

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑