超越基准分数:审计胸部X光结核病筛查的医学视觉-语言模型
Beyond Benchmark Scores: Auditing Medical Vision-Language Models for Chest X-Ray Tuberculosis Screening
浏览论文内容
中文总结 AI 辅助
本研究审计医学视觉-语言模型在胸部X光结核病筛查中的表现,发现模型排名、分数可靠性和阈值保留在不同评估条件下不稳健,强调评估规范完整性的重要性。
中文摘要 AI 辅助
医学模型的基准分数并不能证明在不同评估条件下会得出相同结论。本研究测试了关于模型排名、分数可靠性和筛查性能的主张在队列、提示、阴性谱、指定患病率和操作阈值变化时是否仍然成立。我们审计了三个医学视觉-语言模型(BioMedCLIP、CheXficient和MedSigLIP)以及一个通用领域的OpenCLIP对照模型,使用来自四个数据集(Montgomery、Shenzhen、TBX11K和VinDr-CXR)的12,200条胸部X光影像记录。五个固定的提示族产生了244,000个模型-图像-提示分数。没有模型在所有队列和可靠性标准中均领先。提示族的变化在48个多重性控制的比较中改变了21个的AUROC。将健康对照替换为患病非结核对照后,所有四个模型的AUROC降低了0.075至0.306。在VinDr-CXR上,三个医学模型区分结核病与无发现对照的能力显著优于区分肺炎或肺肿瘤;它们对这两种指定疾病的AUROC点估计均低于0.5。CheXficient有记录的VinDr-CXR预训练暴露,这限制了对其结果的解释。在TBX11K训练集上为95%敏感性选择的阈值,在16个目标评估中仅4个通过点估计保留了该约束。一个五种子监督源模型在TBX11K验证集上达到0.999的AUROC,但在两个外部队列上仅为0.629。保守排除感知重叠候选者缩小了这一差距但未完全消除。这些回顾性、单任务结果表明,区分能力、分数可靠性和阈值保留支持不同的可移植性主张。胸部X光结核病筛查的证据应指明完整的评估规范,而不是将临床可移植性仅归因于检查点。
英文摘要
A medical model's benchmark score does not establish that the same conclusion holds under a different evaluation. This study tests whether claims about model ranking, score reliability and screening performance survive changes in cohort, prompt, negative spectrum, specified prevalence and operating threshold. We audit three medical vision-language models (BioMedCLIP, CheXficient, and MedSigLIP) and a general-domain OpenCLIP comparator on 12,200 chest radiograph records from four datasets (Montgomery, Shenzhen, TBX11K, and VinDr-CXR). Five fixed prompt families yield 244,000 model--image--prompt scores. No model leads every cohort and reliability criterion. Prompt-family changes alter AUROC in 21 of 48 multiplicity-controlled comparisons. Replacing healthy controls with sick non-tuberculosis controls reduces AUROC by 0.075--0.306 across all four models. On VinDr-CXR, the three medical models distinguish tuberculosis from no-finding controls substantially better than from pneumonia or lung tumor; their AUROC point estimates for both named diseases fall below 0.5. CheXficient has documented VinDr-CXR pretraining exposure, which limits the interpretation of its results. Thresholds chosen for 95\% sensitivity on TBX11K training retain that constraint by point estimate in only four of sixteen target evaluations. A five-seed supervised source model reaches 0.999 AUROC on TBX11K validation but 0.629 on each of two external cohorts. Conservative exclusion of perceptual-overlap candidates narrows this gap without closing it. These retrospective, single-task results show that discrimination, score reliability and threshold retention support different portability claims. Evidence for chest X-ray tuberculosis screening should identify the complete evaluation specification rather than attribute clinical portability to a checkpoint alone.
发表机构
- Indian Institute of Technology Indore(印度理工学院印多尔分校)
机构由 AI 辅助整理,请以论文原文为准。