发表机构
Brown University; Brown University Health; Florida International University(布朗大学; 布朗大学健康中心; 佛罗里达国际大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对胸部X射线VLM评估的敏感性,提出可靠性压力测试和决策时路由框架,以改善成本-质量权衡,强调临床部署前需综合考量多种因素。
AI 中文摘要
医学视觉语言模型(VLM)的评估对工作流设计、提示策略和基准构建敏感,但大多数研究将这些因素孤立对待。我们引入了一种基于两个平衡数据集(一个私有报告支持集和一个精选的MIMIC子集)的胸部X射线解读可靠性压力测试。三种医学VLM(CheXagent、MedGemma-4B和MedGemma-27B)在三种提示风格和两种工作流(单VLM和多智能体)下进行评估,产生36种配置。我们表明,仅精确匹配准确率可能高估默认预测为“正常”的保守模型的有效性。诊断可靠性还严重依赖于模型家族和规模:多智能体推理有助于某些配置,但损害其他配置。基于这些观察,我们提出了一种决策时路由框架,仅在有益时选择性地升级到多智能体推理,从而改善固定工作流上的成本-质量权衡。我们的结果强调了在临床部署前,评估协议需要联合考虑提示敏感性、失败模式多样性和工作流选择。
英文摘要
Medical vision-language model (VLM) evaluation is sensitive to workflow design, prompting strategy, and benchmark construction, yet most studies treat these factors in isolation. We introduce a reliability stress test for chest X-ray interpretation built on two balanced datasets (a private report-backed set and a curated MIMIC subset). Three medical VLMs (CheXagent, MedGemma-4B, and MedGemma-27B) are evaluated across three prompt styles and two workflows (single-VLM and multi-agent), producing 36 configurations. We show that exact-match accuracy alone can overstate the effectiveness of conservative models that default to "Normal" predictions. Diagnostic reliability also depends heavily on model family and scale: multi-agent reasoning helps some configurations but hurts others. Building on these observations, we propose a decision-time routing framework that selectively escalates to multi-agent inference only when beneficial, improving the cost-quality trade-off over fixed workflows. Our results highlight the need for evaluation protocols that jointly consider prompt sensitivity, failure-mode diversity, and workflow choice before clinical deployment.
Comments6 pages, 4 figures, 5 tables. Accepted version of a workshop paper presented orally at the 2026 IEEE/ACM Conference on Connected Health: Applications, Systems and Engineering Technologies (CHASE). Code and per-case MIMIC-CXR results: https://github.com/Yangxinyee/cxr-vlm-routing
Journal ref2026 IEEE/ACM Conference on Connected Health: Applications, Systems and Engineering Technologies (CHASE), pp. 428-433
DOI:10.1109/CHASE69719.2026.00073