面向临床医生社交媒体的医学图像识别视觉大语言基础模型
A visual large language foundational model for medical image recognition using clinician-contributed online resources
AI总结:
本研究利用临床医生社交媒体构建ThoughtMed-1M医学VQA数据集,并训练FOLTMed模型,在42个基准上取得85.4%宏准确率,性能领先3-5%。
AI中文摘要:
大语言模型(LLMs)已在多个领域展现出强大的能力,在医学方面显示出相当大的潜力。然而,由于缺乏能够捕捉临床推理和显式图像-文本对齐的视觉问答(VQA)数据集,它们在医学环境中的应用仍然受限。在此,我们利用临床医生导向社交媒体上共享的去标识化医学图像和专家评论。通过将先进的大语言模型与临床医生参与验证相结合,我们建立了一个严格的流程来构建ThoughtMed-1M,这是一个包含超过一百万对VQA的长篇医学VQA数据集,旨在捕捉结构化的临床逻辑和医学图像-文本对齐。为展示其效用,我们开发了基于ThoughtMed-1M训练的基础大语言模型(FOLTMed)。FOLTMed在42个医学VQA基准数据集上取得了最先进的性能,宏准确率达到85.4%,并在ThoughtMed-1M测试集上生成了更具临床连贯性的回答。它在事实性和相似性指标上比最先进的模型高出3%至5%,凸显了一种可扩展的范式,用于推进临床基础多模态大语言模型的研究。
英文摘要:
Large language models (LLMs) have demonstrated strong capabilities across diverse domains, showing considerable potential in medicine. However, their application in medical settings remains limited by the scarcity of visual question answering (VQA) datasets that capture clinical reasoning and explicit image-text alignment. Here, we leverage de-identified medical images and expert commentaries shared through clinician-oriented online resources. By combining an advanced LLM with clinician-in-the-loop verification, we established a rigorous pipeline to construct ThoughtMed-1M, a long-form medical VQA dataset containing over one million VQA pairs and designed to capture structured clinical reasoning and medical image-text alignment. To demonstrate its utility, we developed a FOundational LLM Trained on ThoughtMed-1M (FOLTMed). FOLTMed achieved state-of-the-art performance across 42 medical VQA benchmark datasets, with a macro accuracy of 85.4 percent. It also generated more clinically coherent responses on the ThoughtMed-1M test set, outperforming state-of-the-art models by 3 to 5 percent across factuality and similarity metrics, highlighting a scalable paradigm for advancing research on clinically grounded multimodal LLMs.