MedUAG:面向医学多模态模型的统一理解与生成框架
MedUAG: Unified Understanding and Generation for Medical Multimodal Models
浏览论文内容
中文总结 AI 辅助
针对医学多模态领域缺乏统一理解生成框架的问题,本文构建了MedUAGCorpus数据集与MedUAGBench基准,开发了MedUAG模型,其在医学理解与生成任务中表现出色,为下一代医学多模态系统奠定基础。
中文摘要 AI 辅助
近期,多模态大语言模型(MLLMs)正快速发展为统一理解与生成(UAG)框架。然而,将这类统一范式扩展至医学领域面临两大阻碍:一是缺乏全面的训练与评估基准,二是缺少经过广泛验证的统一医学模型。为解决这些缺口,本文构建了医学UAG的全面基础:首先,打造了迄今为止最大的统一医学理解与生成数据集MedUAGCorpus,涵盖14种成像模态的超600万条实例;其次,推出了系统基准MedUAGBench,该基准将医学生成评估扩展至12项多样化任务,并采用标准化协议;最后,借助上述资源,开发了端到端训练的统一医学模型MedUAG。大量实验表明,MedUAG在广泛的理解与生成任务中表现出色,建立了具有竞争力的基线,为下一代医学多模态系统铺平了道路。
英文摘要
Recent Multimodal Large Language Models (MLLMs) are rapidly evolving into unified understanding and generation (UAG) frameworks. However, extending these unified paradigms to the medical domain is hindered by: the absence of comprehensive training and evaluation benchmarks, and the lack of broadly validated unified medical model. To address these gaps, we present a comprehensive foundation for medical UAG. First, we construct MedUAGCorpus, the largest unified medical understanding and generation dataset to date, comprising over 6 million instances across 14 imaging modalities. Second, we introduce MedUAGBench, a systematic benchmark that expands medical generation evaluation to 12 diverse tasks under standardized protocols. Finally, leveraging these resources, we develop MedUAG, an end-to-end trained unified medical model. Extensive experiments demonstrate that MedUAG achieves strong performance across a wide array of understanding and generation tasks, establishing a competitive baseline and paving the way for next-generation medical multimodal systems.
发表机构
- Zhejiang University(浙江大学)
- Hong Kong University of Science and Technology(香港科技大学)
- Tsinghua University(清华大学)
- Tencent Jarvis Lab(腾讯Jarvis实验室)
机构由 AI 辅助整理,请以论文原文为准。