arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2606.20723cs.CV

医学视觉语言模型 HuluMed 和 MedGemma 以及通用聊天机器人 Gemma 3、ChatGPT Plus 和 Claude Pro 在真实未见伤口图像上的评估

Evaluation of Medical Vision Language Models HuluMed and MedGemma, and general purpose chatbots Gemma 3, ChatGPT Plus, and Claude Pro on real previously unseen wound images

  • Department of Computer Science, New Jersey Institute of Technology(新泽西理工学院计算机科学系)
  • Vascular and Endovascular Surgery, Robert Wood Johnson Hospital(罗伯特·伍德·约翰逊医院血管外科)
  • Department of Data Science, New Jersey Institute of Technology(新泽西理工学院数据科学系)

机构由 AI 辅助整理,请以论文原文为准。

Yunzhe Xue, Mohammed Saim Ahmed Quadri, Neal Panse, Justin W. Ady, Usman Roshan

AI总结:

本研究评估了六种视觉语言模型在慢性伤口分析任务上的表现,发现通用模型 ChatGPT 和 Claude 显著优于医学专用模型,表明广泛的多模态推理能力比领域知识更重要。

AI中文摘要:

慢性伤口评估仍然是一项临床挑战性任务,需要准确解读伤口形态、组织成分、血管特征和感染风险。视觉语言模型(VLM)的最新进展通过图像理解与临床推理相结合,引入了自动化多模态伤口分析的可能性。本研究使用一个扩展的、精心整理的包含20个临床多样化伤口的数据集(涵盖血管、手术、缺血、静脉、淋巴水肿和截肢相关病因),评估了多个通用和医学专用的开源及专有VLM在临床伤口评估中的性能。使用一个包含12个问题的结构化临床框架对六个VLM进行了评估,涵盖伤口分类、感染风险、血管干预建议、清创紧迫性、伤口治疗选择和高级管理规划。在20个伤口病例和240个临床医生评级的伤口分析决策中,ChatGPT以174/240(72.50%)的正确率获得最高总体性能,其次是Claude的149/240(62.08%)。在开源和医学专用模型中,HuluMed以96/240(40.00%)的正确率表现最强,其次是Gemma 3(81/240,33.75%)、MedGemma 4B(62/240,25.83%)和MedGemma 27B(42/240,17.50%)。研究结果表明,前沿通用多模态系统目前展现出比医学专用替代方案明显更强的伤口分析性能,凸显了广泛多模态推理能力与领域特定医学知识并重的重要性。尽管当前VLM在临床决策支持方面展现出有前景的潜力,但在高级伤口管理推理、程序规划和自主临床可靠性方面仍存在显著局限性。

英文摘要:

Chronic wound assessment remains a clinically challenging task that requires accurate interpretation of wound morphology, tissue composition, vascular characteristics, and infection risk. Recent advances in Vision-Language Models (VLMs) have introduced the possibility of automated multimodal wound analysis through image understanding combined with clinical reasoning. This study evaluates the performance of several general-purpose and medically specialized open-source and proprietary VLMs for clinical wound assessment using an expanded, curated dataset of 20 clinically diverse wounds spanning vascular, surgical, ischemic, venous, lymphedema, and amputation-related etiologies. Six VLMs were evaluated using a structured twelve-question clinical framework covering wound classification, infection risk, vascular intervention recommendations, debridement urgency, wound therapy selection, and advanced management planning. Across 20 wound cases and 240 clinician-graded wound-analysis decisions, ChatGPT achieved the highest overall performance with 174/240 correct responses (72.50%), followed by Claude with 149/240 (62.08%). Among the open-source and medically specialized models, HuluMed achieved the strongest performance with 96/240 correct responses (40.00%), followed by Gemma 3 (81/240, 33.75%), MedGemma 4B (62/240, 25.83%), and MedGemma 27B (42/240, 17.50%). The findings suggest that frontier general-purpose multimodal systems currently demonstrate substantially stronger wound-analysis performance than medically specialized alternatives, highlighting the continued importance of broad multimodal reasoning capabilities alongside domain-specific medical knowledge. Although current VLMs demonstrate promising potential for clinical decision support, substantial limitations remain in advanced wound-management reasoning, procedural planning, and autonomous clinical reliability.

↑