MVC-Bench:医学视觉-语言模型的校准基准测试
MVC-Bench: Benchmarking Calibration of Medical Vision-Language Models
浏览论文内容
中文总结 AI 辅助
该研究针对医学视觉-语言模型校准不足的问题,提出以校准为核心的基准MVC-Bench,通过多维度评估并提出MCM正则化方法,为医疗工作流提供校准改进指导。
中文摘要 AI 辅助
对视觉-语言模型(VLMs)和医学视觉-语言模型(Medical-VLMs)的可靠评估需要校准后的置信度,尤其是在现实临床场景下。然而,现有研究主要聚焦于提升准确率,却忽视了医学领域的校准问题。为此,我们提出MVC-Bench,这是一个面向医学图像分类、以校准为核心的基准测试,适用于VLMs和Medical-VLMs。MVC-Bench从三个维度评估校准性能:(i)对模态、主干网络及域偏移的鲁棒性;(ii)校准策略与提示微调方法的有效性;(iii)在提示模板和随机种子变化下的稳定性。该基准测试涵盖8种不同主干网络、3种医学模态(包括眼底成像、组织病理学和胸部X射线),并设置了域内和域偏移两种场景。它对比了事后校准、训练时校准和零样本推理方法,以及6种提示微调方法。在超过1638项受控实验中,我们以准确率和预期校准误差(ECE)作为核心指标,同时补充报告了最大校准误差(MCE)和自适应校准误差(ACE)等校准指标的结果。我们进一步探究了VLMs和Medical-VLMs校准误差的潜在原因,并提出一种简单的训练时校准方法——多类边际(Multi-Class Margin, MCM)正则化,该方法在域内场景的12个设置中,有10个达到了最低ECE,且在域偏移场景下仍具竞争力。总体而言,MVC-Bench为安全关键的医疗工作流提供了结构化评估框架和可操作的校准改进指导。
英文摘要
Reliable evaluation of vision-language models (VLMs) and medical vision-language models (Medical-VLMs) requires calibrated confidence, particularly under realistic clinical conditions. However, existing efforts mainly focused on improving accuracy, leaving calibration in the medical domain underexplored. To this end, we propose MVC-Bench, a calibration-centric benchmark for medical image classification with VLMs and Medical-VLMs. MVC-Bench assesses the calibration across three axes: (i) robustness to modality, backbone, and domain shift (ii) effectiveness of calibration strategies and prompt-tuning methods (iii) stability under prompt-template and random-seed variations. The benchmark covers eight different backbones, three medical modalities, including fundus imaging, histopathology, and chest X-ray under in-domain and domain shift settings. It compares post-hoc calibration, train-time calibration, and zero-shot inference methods, together with six prompt-tuning methods. Across more than 1638 controlled experiments, we report accuracy and Expected Calibration Error (ECE) as primary metrics, and further report results with complementary calibration measures, including Maximum Calibration Error (MCE) and Adaptive Calibration Error (ACE). We further investigate the underlying causes of miscalibration in VLMs and Medical-VLMs and propose a simple train-time calibration method, Multi-Class Margin (MCM) regularization, which achieves lowest ECE on 10 out of 12 settings in in-domain and remains competitive under domain shifts. Collectively, MVC-Bench provides a structured evaluation framework and actionable guidance for improving calibration in safety-critical medical workflows.
发表机构
- Mohamed bin Zayed University of AI(穆罕默德·本·扎耶德人工智能大学)
- Sabaragamuwa University of Sri Lanka(斯里兰卡萨巴拉加穆瓦大学)
- Digital Platform Development, SLT PLC(SLT公共有限公司数字平台开发部)
机构由 AI 辅助整理,请以论文原文为准。