arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2507.17539cs.AIcs.CVeess.IV

通过临床认知链推理构建用于定位-诊断协作的眼科多模态大语言模型

Constructing Ophthalmic MLLM for Positioning-diagnosis Collaboration Through Clinical Cognitive Chain Reasoning

  • Shanghai Artificial Intelligence Laboratory(上海人工智能实验室)
  • University of Science and Technology of China(中国科学技术大学)

机构由 AI 辅助整理,请以论文原文为准。

Xinyao Liu, Diping Song

更新

AI总结:

提出眼科专用 MLLM FundusExpert 及 FundusGen 数据集,通过定位-诊断认知链提升眼底问答与零样本报告生成性能。

AI中文摘要:

多模态大语言模型(MLLMs)在医学诊断领域展现出显著潜力。然而,它们在眼科等专科领域面临关键挑战,尤其是标注粒度碎片化和临床推理逻辑不一致,这阻碍了精确的跨模态理解。本文介绍了 FundusExpert,这是一种眼科专用 MLLM,具备集成的定位-诊断推理能力;同时介绍了 FundusGen,这是一个通过智能 Fundus-Engine 系统构建的数据集。Fundus-Engine 可自动完成定位,并利用基于 MLLM 的语义扩展,在单张眼底图像内整合全局疾病分类、局部目标检测和细粒度特征分析。此外,通过构建与临床一致的认知链,它引导模型生成可解释的推理路径。FundusExpert 使用来自 FundusGen 的指令数据进行微调,在眼科问答任务中取得最佳性能,其平均准确率比 40B MedRegA 高出 26.6%。它在零样本报告生成任务中也表现出色,达到 77.0% 的临床一致性,显著优于 GPT-4o 的 47.6%。此外,我们揭示了数据质量与模型能力之间的缩放定律($L \propto N^{0.068}$),证明 FundusGen 中的认知对齐标注提升了数据利用效率。通过将区域级定位与诊断推理链相结合,本工作开发了一种可扩展、与临床对齐的 MLLM,并探索了一条弥合特定 MLLM 中视觉-语言鸿沟的路径。我们的项目可在 https://github.com/MeteorElf/FundusExpert 获取。

英文摘要:

Multimodal large language models (MLLMs) demonstrate significant potential in the field of medical diagnosis. However, they face critical challenges in specialized domains such as ophthalmology, particularly the fragmentation of annotation granularity and inconsistencies in clinical reasoning logic, which hinder precise cross-modal understanding. This paper introduces FundusExpert, an ophthalmology-specific MLLM with integrated positioning-diagnosis reasoning capabilities, along with FundusGen, a dataset constructed through the intelligent Fundus-Engine system. Fundus-Engine automates localization and leverages MLLM-based semantic expansion to integrate global disease classification, local object detection, and fine-grained feature analysis within a single fundus image. Additionally, by constructing a clinically aligned cognitive chain, it guides the model to generate interpretable reasoning paths. FundusExpert, fine-tuned with instruction data from FundusGen, achieves the best performance in ophthalmic question-answering tasks, surpassing the average accuracy of the 40B MedRegA by 26.6%. It also excels in zero-shot report generation tasks, achieving a clinical consistency of 77.0%, significantly outperforming GPT-4o's 47.6%. Furthermore, we reveal a scaling law between data quality and model capability ($L \propto N^{0.068}$), demonstrating that the cognitive alignment annotations in FundusGen enhance data utilization efficiency. By integrating region-level localization with diagnostic reasoning chains, our work develops a scalable, clinically-aligned MLLM and explores a pathway toward bridging the visual-language gap in specific MLLMs. Our project can be found at https://github.com/MeteorElf/FundusExpert.

↑