发表机构
Shandong University; Stanford University(山东大学; 斯坦福大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究视觉语言模型应用于医学AI时原始医学数据标准化问题,通过构建基准测试,让模型处理原始数据集文件夹,评估其多种能力,发现即使最佳模型端到端成功率也低,凸显该环节是医学AI诊断关键瓶颈。
AI 中文摘要
随着视觉语言模型(VLMs)越来越多地应用于医学人工智能,现有基准主要关注评估其对给定医学图像和文本的诊断能力,隐含地假设标准化的医学图像、文本或问答对已经准备好。然而,当我们在实际临床实践中应用VLMs时,这一假设并不成立,因为医学数据通常是原始的、异构的,并且分散在不同来源。在本文中,我们研究了这一缺失的步骤,即原始医学数据标准化。具体来说,模型被给予原始数据集文件夹,并评估它们识别源格式、将原始医学图像转换为与VLM兼容的视觉输入、提取相关文本信息以及将结果组织成结构化图像-文本对的能力。为了构建这个医学数据标准化基准(MDS-Bench),我们手动注释了1939个原始医学数据标准化任务,涵盖了不同的临床实践、放射学模态、注释格式和目录布局。广泛的实验表明,即使是性能最好的VLMs,即Gemini 3 Flash,端到端成功率也仅为48.6%。我们的研究强调了原始医学数据标准化是实际医学人工智能诊断的关键瓶颈。
英文摘要
As vision-language models (VLMs) are increasingly applied to medical AI, existing benchmarks mainly focus on evaluating their diagnostic ability over given medical images and texts, implicitly assuming that standardized medical images, texts, or question-answer pairs are already prepared. However, this assumption does not hold when we apply VLMs in real clinical practice, where medical data is often raw, heterogeneous, and fragmented across different sources. In this paper, we study this missing step, i.e., raw medical data standardization. Specifically, models are given raw dataset folders and evaluated on their ability to identify source formats, convert raw medical images into VLM-compatible visual inputs, extract relevant textual information, and organize the results into structured image-text pairs. To construct this Medical Data Standardization Benchmark (MDS-Bench), we manually annotate 1,939 raw medical data standardization tasks covering diverse clinical practice, radiology modalities, annotation formats, and directory layouts. Extensive experiments show that even the best performing VLM, i.e., Gemini 3 Flash, achieves only a 48.6% end-to-end success rate. Our research highlights raw medical data standardization as a critical bottleneck for medical AI diagnosis in real practice.
CommentsAccepted to EMNLP 2026 Main Conference