RetiBridge:用知识引导的多模态大语言模型连接定量视网膜生物标志物与定性诊断
RetiBridge: Bridging Quantitative Retinal Biomarkers and Qualitative Diagnosis with a Knowledge-Guided Multimodal Large Language Model
- University of Liverpool(利物浦大学)
- Institute of Life Course & Medical Sciences(生命课程与医学科学研究院)
- Department of Eye and Vision Sciences(眼科与视觉科学系)
- Computer Science Department(计算机科学系)
- Cardiovascular & Metabolic Medicine(心血管与代谢医学)
- Liverpool Centre for Cardiovascular Science(利物浦心血管科学中心)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
该研究提出RetiBridge模型,结合CFP、OCT数据与文本,通过知识引导微调学习定量到定性诊断路径,在眼科理解基准上表现优于开源基线及OpenAI o3,代码数据已公开。
AI中文摘要:
彩色眼底照相(CFP)和光学相干断层扫描(OCT)捕获的视网膜生物标志物为眼部及全身疾病提供了具有临床价值的证据。多模态大语言模型(MLLM)在视网膜图像解读方面展现出潜力,但现有眼科模型很少对这些临床相关生物标志物进行量化,也很少将其测量结果明确转化为基于证据的定性诊断结论。为解决这一差距,我们提出RetiBridge,这是一种知识引导的多模态大语言模型,可联合分析CFP、OCT和文本,明确将定量视网膜生物标志物与定性临床子推理及连贯诊断结论连接起来。RetiBridge结合了知识引导的指令生成、OCT-生物标志物对齐以及监督多模态指令微调,以学习基于生物标志物的定量到定性诊断路径。我们使用英国生物银行(UK Biobank)的15611对CFP-OCT样本(包含31种OCT生物标志物和6种CFP生物标志物)构建了基于证据的眼科理解基准(Grounded Ophthalmic Understanding benchmark),用于评估诊断分类、报告生成质量和细粒度临床质量。尽管仅对70亿参数的Qwen2主干进行了基于LoRA的微调,RetiBridge仍优于所有评估的开源70亿和320亿参数基线模型,在定量准确率、证据依据、覆盖完整性和BERTScore方面取得了最高成绩,在这些关键的基于生物标志物的指标上甚至超过了OpenAI o3。我们的代码和数据已在RetiBridge代码库中发布。
英文摘要:
Retinal biomarkers captured by color fundus photography and optical coherence tomography provide clinically valuable evidence for both ocular and systemic diseases. Multimodal large language models (MLLMs) have shown promise for retinal image interpretation, yet existing ophthalmic models rarely quantify these clinically relevant biomarkers or explicitly translate their measurements into qualitative, evidence-grounded diagnostic conclusions. To address this gap, we introduce RetiBridge, a knowledge-guided multimodal large language model that jointly analyzes color fundus photography (CFP), optical coherence tomography (OCT), and text, explicitly bridging quantitative retinal biomarkers to qualitative clinical sub-inferences and coherent diagnostic conclusions. RetiBridge combines knowledge-guided instruction generation, OCT-biomarker alignment, and supervised multimodal instruction tuning to learn a biomarker-grounded quantitative-to-qualitative diagnostic pathway. Using 15,611 paired CFP-OCT samples from UK Biobank with 31 OCT and 6 CFP biomarkers, we construct the Grounded Ophthalmic Understanding benchmark to evaluate diagnostic classification, report generation quality, and fine-grained clinical quality. Despite using only LoRA-based fine-tuning of a 7B-parameter Qwen2 backbone, RetiBridge outperforms all evaluated open-source 7B and 32B baselines, achieving the highest quantitative accuracy, evidence grounding, coverage completeness, and BERTScore, while surpassing OpenAI o3 on these key biomarker-grounded metrics. Our code and data are released in the RetiBridge repository.