缩小眼科人工智能的差距:MM-Retinal-Reason数据集与OphthaReason模型用于动态多模态推理
Bridging the Gap in Ophthalmic AI: MM-Retinal-Reason Dataset and OphthaReason Model toward Dynamic Multimodal Reasoning
- School of Computer Science and Engineering, Southeast University(东南大学计算机科学与工程学院)
- Department of Ophthalmology, The First Affiliated Hospital of Nanjing Medical University(南京医科大学第一附属医院眼科学系)
- School of Computer Science and Engineering, Nanjing University of Science and Technology(南京理工大学计算机科学与工程学院)
- School of Computer Science, Northwestern Polytechnical University(西北工业大学计算机学院)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
该研究针对现有眼科AI仅聚焦基础推理的问题,构建首个眼科多模态数据集MM-Retinal-Reason,并提出带UADT方法的OphthaReason模型,实现眼科多模态推理性能的显著提升。
AI中文摘要:
多模态大语言模型(MLLMs)近期在强化学习范式下展现出出色的推理能力。尽管医疗领域已探索出若干多模态推理模型,但多数仅聚焦于基础推理,即基于视觉特征匹配的浅层推断。然而,现实临床诊断远超基础推理范畴,需整合异质性临床信息(如主诉与病史)及多模态医学影像数据的推理过程。为缩小该差距,我们推出MM-Retinal-Reason,这是首个涵盖感知与推理全流程的眼科多模态数据集,包含基础推理与复杂推理任务,旨在提升以视觉为核心的基础推理能力并模拟真实临床思维模式。基于MM-Retinal-Reason,我们提出OphthaReason,首个具备分步推理轨迹的眼科专用多模态推理模型。为灵活适配基础与复杂推理任务,我们专门设计了名为不确定性感知动态思维(UADT)的新方法,该方法通过熵估计样本级不确定性,并利用塑形优势机制动态调整模型的探索深度。综合实验表明,我们的模型在基础与复杂推理任务上均实现了最优性能,较通用MLLMs、医疗MLLMs、基于强化学习的医疗MLLMs及眼科MLLMs分别至少提升24.92%、15.00%、21.20%、17.66%。项目主页:link。
英文摘要:
Multimodal large language models (MLLMs) have recently demonstrated remarkable reasoning abilities under reinforcement learning (RL) paradigm. However, most existing multimodal medical reasoning models focus on basic reasoning, which refers to shallow inference based on visual feature matching. In contrast, real-world clinical diagnosis extends beyond basic reasoning, demanding complex reasoning that integrates heterogeneous clinical information (such as chief complaints and medical history) with multimodal medical imaging data. To bridge this gap, we introduce MM-Retinal-Reason, an ophthalmic multimodal dataset covering the full spectrum of perception and reasoning. Specifically, it is the first dataset in ophthalmology to encompass both basic and complex reasoning tasks with Chain-of-Thought (CoT) trajectories, aiming to enhance visual-centric reasoning and emulate realistic clinical decision-making. Building upon MM-Retinal-Reason, we propose OphthaReason, the first RL-enhanced ophthalmic multimodal reasoning model with step-by-step reasoning traces. To enable flexible adaptation to both basic and complex reasoning tasks, we further introduce Uncertainty-Aware Dynamic Thinking (UADT), which estimates sample-level uncertainty via entropy and dynamically modulates exploration depth through a shaped advantage mechanism. Comprehensive experiments demonstrate the effectiveness of our model on both basic and complex reasoning tasks, outperforming general-purpose MLLMs, medical MLLMs, RL-based medical MLLMs, and ophthalmic MLLMs by at least 15.47\%. Project Page: \href{https://github.com/lxirich/OphthaReason}{link}.