arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2604.08322cs.CV

Fundus-R1: 基于公共数据训练具备知识感知推理能力的视网膜图像阅读MLLM

Fundus-R1: Training a Fundus-Reading MLLM with Knowledge-Aware Reasoning on Public Data

  • Renmin University of China(中国人民大学)
  • The Hong Kong University of Science and Technology(香港科技大学)
  • Zhejiang Gongshang University(浙江工商大学)

机构由 AI 辅助整理,请以论文原文为准。

Yuchuan Deng, Qijie Wei, Kaiheng Qian, Jiazhen Liu, Zijie Xin, Bangxiang Lan, Jingyu Liu, Jianfeng Dong, Xirong Li

更新

AI总结:

本文提出Fundus-R1,通过公共数据集训练具备知识感知推理能力的视网膜图像阅读MLLM,采用RAG方法生成知识感知推理轨迹,并通过过程奖励增强RLVR,实验证明其在三个基准测试中优于多个基线模型。

AI中文摘要:

视网膜影像如眼底照相、OCT和UWF对于早期发现视网膜异常和疾病至关重要。由于其知识密集性,视网膜图像理解是一个具有挑战性的视觉-语言任务。一种新兴的方法是通过监督微调(SFT)或强化学习与可验证奖励(RLVR)在大量内部样本上微调通用多模态大语言模型(MLLM)。然而,这些有价值的样本不公开,这不仅阻碍了可重复性,也限制了研究到少数玩家。为克服这一障碍,我们尝试训练一个增强推理的视网膜图像阅读MLLM,称为Fundus-R1,使用仅公共数据集,其中超过94%的数据仅带有图像级别标签。我们的技术贡献是双方面的。首先,我们提出了一种基于RAG的方法,用于生成图像特定、知识感知的推理轨迹。此类自动生成的轨迹将通用MLLM识别的视觉发现与图像标签联系起来,基于眼科知识。第二,我们通过过程奖励增强RLVR,鼓励生成推理轨迹在每次回放中的自一致性。在三个视网膜图像阅读基准测试(即FunBench、Omni-Fundus和GMAI-Fundus)上进行了广泛的实验,结果显示Fundus-R1在多个基线模型中表现明显优于,包括其通用对应物(Qwen2.5-VL)和一个经过强化学习训练但不使用生成轨迹的更强版本。这项工作为利用公开数据训练强大的视网膜图像阅读MLLM铺平了道路。

英文摘要:

Fundus imaging such as CFP, OCT and UWF is crucial for the early detection of retinal anomalies and diseases. Fundus image understanding, due to its knowledge-intensive nature, poses a challenging vision-language task. An emerging approach to addressing the task is to post-train a generic multimodal large language model (MLLM), either by supervised finetuning (SFT) or by reinforcement learning with verifiable rewards (RLVR), on a considerable amount of in-house samples paired with high-quality clinical reports. However, these valuable samples are not publicly accessible, which not only hinders reproducibility but also practically limits research to few players. To overcome the barrier, we make a novel attempt to train a reasoning-enhanced fundus-reading MLLM, which we term Fundus-R1, using exclusively public datasets, wherein over 94\% of the data are annotated with only image-level labels. Our technical contributions are two-fold. First, we propose a RAG-based method for composing image-specific, knowledge-aware reasoning traces. Such auto-generated traces link visual findings identified by a generic MLLM to the image labels in terms of ophthalmic knowledge. Second, we enhance RLVR with a process reward that encourages self-consistency of the generated reasoning trace in each rollout. Extensive experiments on three fundus-reading benchmarks, i.e., FunBench, Omni-Fundus and GMAI-Fundus, show that Fundus-R1 clearly outperforms multiple baselines, including its generic counterpart (Qwen2.5-VL) and a stronger edition post-trained without using the generated traces. This work paves the way for training powerful fundus-reading MLLMs with publicly available data.

↑