发表机构
DAMO Academy, Alibaba Group; Hupan Laboratory; College of Computer Science and Technology, Zhejiang University; Department of Computer Science and Technology, Tsinghua University; Department of Radiology, The Affiliated Yangming Hospital of Ningbo University; Zhejiang University-University of Illinois Urbana-Champaign Institute, Zhejiang University; Hepato-Pancreato-Biliary Center, Beijing Tsinghua Changgung Hospital, School of Clinical Medicine, Tsinghua Medicine, Tsinghua University; School of Software, Tsinghua University; Beijing National Research Center for Information Science and Technology, Tsinghua University(达摩院,阿里巴巴集团; 湖畔实验室; 浙江大学计算机科学与技术学院; 清华大学计算机科学与技术系; 宁波大学附属阳明医院放射科; 浙江大学伊利诺伊大学厄巴纳香槟校区联合学院,浙江大学; 清华长庚医院肝胆胰中心,清华大学临床医学院,清华医学,清华大学; 清华大学软件学院; 清华大学北京信息科学与技术国家研究中心)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究针对多模态大语言模型在医学领域部署的挑战,提出ClinFusion,采用组合级联视觉编码器架构和视觉基础评估框架,在多模态医学基准测试中表现优异,超越开源和专有模型,经专家盲评验证效果良好。
AI 中文摘要
多模态大语言模型在革新临床实践方面潜力巨大,但在医学领域部署存在以视觉为中心的挑战。本文介绍ClinFusion,它是专为整体医学理解设计的以视觉为中心的多模态大语言模型。提出组合式和级联式视觉编码器架构,含级联空间感知局部融合算子。还引入视觉基础评估框架。ClinFusion在多模态医学基准测试中表现出色,超越领先开源模型,在部分基准测试中优于专有模型,还可通过智能工具用于临床工作流程。经放射科医生盲评,ClinFusion生成的报告排名最高,验证了基于感兴趣区域的指标与专家判断相关性最强。
英文摘要
Multimodal large language models (MLLMs) hold immense potential to revolutionize clinical practice, yet deploying them in the medical domain is fundamentally a vision-centric challenge: models must absorb knowledge from heterogeneous 2D and 3D medical images, and evaluation protocols must align with radiologists' clinical practice and provide an accurate, fine-grained and factualness-driven assessment. In this paper, we introduce ClinFusion, a vision-centric MLLM designed for holistic medical understanding that systematically addresses these limitations. We propose a compositional and cascaded vision encoder architecture featuring a Cascade Spatial-Aware Locality Fusion operator that unifies diverse 2D and native 3D medical image understanding within a fused encoder. We further introduce a vision-grounded evaluation framework, including MedIF-Bench for instruction-following assessment and a region-of-interest-grounded method for clinically aligned and factualness-driven report generation evaluation. We show that ClinFusion sets a new state-of-the-art across a comprehensive suite of 2D and 3D multimodal medical benchmarks---spanning visual question answering, report generation, and instruction following---as well as textual medical tasks, outperforming leading open-source medical MLLMs (\textit{e.g.}, Hulu-Med, Lingshu) on 20 out of 24 benchmarks and demonstrating multimodal capabilities better than powerful proprietary models such as GPT-5.2 and Gemini-3-Flash on 13 out of 16 benchmarks, and can be further augmented with agentic tool use for retrieval-augmented and tool-assisted clinical workflows. A blinded evaluation by board-certified radiologists confirms that ClinFusion produces the highest-ranked reports, and validates our RoI-grounded metric as achieving the strongest correlation with expert judgment among all automatic evaluation metrics examined.
CommentsCode: https://github.com/alibaba-damo-academy/ClinFusion Models: https://huggingface.co/collections/Alibaba-DAMO-Academy/clinfusion