CRS-Bench:面向医学图像编码器的参考相对可靠性基准
CRS-Bench: A Reference-Relative Reliability Benchmark for Medical Image Encoders
- Vanderbilt University Medical Center(范德堡大学医学中心)
- Vanderbilt University(范德堡大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本研究提出CRS-Bench基准,通过评估15个预训练医学图像编码器的多轴可靠性,结合临床可靠性评分,发现部分编码器排序与AUROC结果反转,确定PanDerm等为稳定领先层级,为医学编码器选择提供更全面框架。
AI中文摘要:
预训练图像编码器是医学图像分类的核心,而专家标注成本高昂,任务特定队列往往有限。随着模型空间从通用编码器扩展到广谱医学及专科特定编码器,表征选择成为重要的建模决策。仅干净测试判别力不足以满足此需求:具有相似AUROC的编码器在校准、标签效率以及在采集扰动或分布偏移下的稳定性方面可能存在差异。我们引入CRS-Bench,这是一个用于多目标医学编码器选择的受控基准。CRS-Bench使用ISIC 2019、APTOS 2019和CheXpert数据集,在皮肤病学、眼科学和放射学领域评估15个预训练编码器家族,并以CheXpert到MIMIC-CXR作为观察到的机构偏移,生成17575条受控运行记录和3515条种子聚合指标行。每个编码器沿四个操作可靠性维度表征:判别力、校准、标签效率和鲁棒性。我们使用临床可靠性评分(Clinical Reliability Score, CRS)总结这些维度,该评分是一种帕累托感知、参考相对的评分,结合了优势度、轮廓平衡和最差轴性能。AUROC与CRS呈正相关,但决策不等价:105个成对排序中有21个反转,平均绝对秩位移为1.87。配对种子自举分析确定PanDerm、MedSigLIP和MedGemma为稳定的领先可靠性层级,而非统计上明确的单一领导者。CRS-Bench提供了一个受控框架,用于从多轴可靠性轮廓中选择医学图像编码器,而非仅基于干净测试AUROC。
英文摘要:
Pretrained image encoders are central to medical image classification, where expert annotation is costly and task-specific cohorts are often limited. As the model space expands from general-purpose to broad-medical and specialty-specific encoders, selecting the representation becomes a substantive modeling decision. Clean-test discrimination alone is insufficient for this purpose: encoders with similar AUROC can differ in calibration, label efficiency, and stability under acquisition perturbations or distribution shift. We introduce CRS-Bench, a controlled benchmark for multi-objective medical encoder selection. CRS-Bench evaluates 15 pretrained encoder families across dermatology, ophthalmology, and radiology using ISIC 2019, APTOS 2019, and CheXpert, with CheXpert-to-MIMIC-CXR as an observed institutional shift, yielding 17,575 controlled run records and 3,515 seed-aggregated metric rows. Each encoder is characterized along four operational reliability dimensions: discrimination, calibration, label efficiency, and robustness. We summarize these dimensions using the Clinical Reliability Score (CRS), a Pareto-aware, reference-relative score combining dominance, profile balance, and worst-axis performance. AUROC and CRS are positively associated but not decision-equivalent: 21 of 105 pairwise orderings reverse, with a mean absolute rank displacement of 1.87. Paired-seed bootstrap analysis identifies PanDerm, MedSigLIP, and MedGemma as a stable leading reliability tier rather than a statistically resolved single leader. CRS-Bench provides a controlled framework for selecting medical image encoders from multi-axis reliability profiles rather than clean-test AUROC alone.