MedTVT-R1:赋能医学推理与诊断的多模态大语言模型
MedTVT-R1: A Multimodal LLM Empowering Medical Reasoning and Diagnosis
- The Hong Kong University of Science & Technology (Guangzhou)(香港科技大学(广州))
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
MedTVT-R1提出一种整合临床多模态数据、采用模态感知层和GRPO强化微调的多模态大语言模型,实现可解释的多疾病诊断,并在实验中展现优越性能。
AI中文摘要:
准确且可解释的多疾病诊断仍是医学研究中的关键挑战,尤其是在利用异构多模态医学数据时。现有方法往往依赖单模态数据,限制了其全面理解复杂疾病的能力。为解决这一问题,我们提出了MedTVT-R1,一种新颖的多模态大语言模型(MLLM)框架,旨在整合临床多模态数据进行推理和多疾病诊断。我们构建了MedTVT-QA,一个精选的指令数据集,通过证据链(Chain of Evidence)方法提供生理层面解释和疾病层面诊断的问答对。MedTVT-R1整合了模态感知层,以捕获模态间依赖关系并自适应地加权模态贡献。此外,我们采用基于组相对策略优化(GRPO)的强化微调,并结合Jaccard奖励函数来增强诊断推理能力。实验结果表明,MedTVT-R1在多模态特征利用和多疾病诊断方面具有优越性,为诊断报告生成和合并症推理等临床应用提供了巨大潜力。数据集和代码可在https://github.com/keke-nice/MedTVT-R1获取。
英文摘要:
Accurate and interpretable multi-disease diagnosis remains a critical challenge in medical research, particularly when leveraging heterogeneous multimodal medical data. Current approaches often rely on single-modal data, limiting their ability to comprehensively understand complex diseases. To address this, we propose MedTVT-R1, a novel Multimodal Large Language Model (MLLM) framework designed to integrate clinical multimodal data for reasoning and diagnosing multiple diseases. We construct MedTVT-QA, a curated instruction dataset that provides question-answer pairs for physiological-level interpretations and disease-level diagnoses with a Chain of Evidence approach. MedTVT-R1 incorporates a modality perception layer to capture inter-modal dependencies and adaptively weight modality contributions. Additionally, we employ Group Relative Policy Optimization (GRPO)-based Reinforcement Fine-Tuning with a Jaccard Reward function to enhance diagnostic reasoning. Experimental results demonstrate MedTVT-R1's superiority in multimodal feature utilization and multi-disease diagnosis, offering significant potential for clinical applications such as diagnostic report generation and comorbidity reasoning. The dataset and code are available at https://github.com/keke-nice/MedTVT-R1.