发表机构
Georgia State University; Georgia Institute of Technology; Emory University; TReNDS Center(佐治亚州立大学; 佐治亚理工学院; 埃默里大学; TReNDS中心)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究对6个指令微调VLMs开展脑MRI行为安全审计,发现其置信度校准差、存在大量高置信度错误,提出医学图像VLM评估需报告置信度可靠性等指标。
AI 中文摘要
包括医疗专家模型在内的视觉语言模型(VLMs)正越来越多地被提出用于医学成像任务,然而它们的声明置信度却很少与模型正确性分开评估。本研究以脑MRI作为受控的高风险测试平台,探究前沿多模态系统中更广泛的失效模式:模型可能看似具备能力,却缺乏可靠的自我认知。我们对6个经过指令微调的VLMs(5个通用模型和1个医疗专家模型)开展自动分级的行为审计与试点研究,使用的4102张图像包括250名受试者的4032张轴位/冠状位/矢状位MRI切片,以及70张非脑/噪声对照图像,标签源自公开元数据和已发布的专家分割掩码,而非新的人工标注。所有模型的答案覆盖率接近100%,但口头表述的置信度校准效果很差:预期校准误差(ECE)范围为0.27至0.40,错误答案的平均置信度范围为0.82至0.97,且33%至46%的已回答项目存在高置信度错误。最准确的模型在其错误上也最为自信,而基础模型/专家模型的对比表明,医疗适配提升了肿瘤存在检测的能力,却未改善置信度可靠性。开放式诊断进一步显示,幻觉和弃权(不执行)与多项选择题的准确性各自独立变化。这些发现表明,医学图像VLM评估应在准确性之外,同时报告口头表述的置信度可靠性、置信错误、幻觉和弃权(不执行)情况。
英文摘要
Vision-language models (VLMs), including medical specialists, are increasingly proposed for medical imaging, yet their stated confidence is rarely evaluated separately from correctness. We use brain MRI as a controlled, high-stakes testbed for a broader failure mode in frontier multimodal systems: models can appear competent while lacking reliable self-knowledge. We present an automatically graded behavioral audit and pilot study of six instruction-tuned VLMs (five general-purpose and one medical specialist) on 4,102 images (4,032 axial/coronal/sagittal MRI slices from 250 subjects plus 70 non-brain/noise controls), with labels derived from public metadata and released expert segmentation masks rather than new human annotation. Across models, answer coverage is near-complete, but verbalized-confidence calibration is poor: ECE ranges from 0.27 to 0.40, mean confidence on incorrect answers ranges from 0.82 to 0.97, and 33-46% of answered items are high-confidence errors. The most accurate model is also the most confident on its errors, while a base/specialist family contrast suggests that medical adaptation improves tumor-presence detection without improving confidence reliability. Open-ended diagnostics further show that hallucination and abstention vary separately from multiple-choice accuracy. These findings argue that medical-image VLM evaluation should report verbalized-confidence reliability, confident error, hallucination, and abstention alongside accuracy.