arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2511.22018cs.CVcs.AI

MedEyes: 学习动态视觉聚焦以进行医学逐步诊断

MedEyes: Learning Dynamic Visual Focus for Medical Progressive Diagnosis

Chunzheng Zhu, Yangfang Lin, Shen Chen, Yijun Wang, Jianxin Lin

更新

AI总结:

MedEyes通过动态视觉聚焦和双模式探索策略,提升医学逐步诊断的准确性与临床相关性。

AI中文摘要:

准确的医学诊断往往涉及逐步的视觉聚焦和迭代推理,这是临床工作流程中常见的特征。尽管最近的视觉-语言模型通过可验证奖励的强化学习(RLVR)展示了有希望的链式推理(CoT)能力,但其纯粹的策略学习范式倾向于强化表面上连贯但临床上不准确的推理路径。我们提出了MedEyes,一种新的强化学习框架,通过逐步关注和解释相关的医学图像区域,动态建模医生风格的诊断推理。通过将离策略专家指导纳入其中,MedEyes将专家的视觉搜索轨迹转换为结构化的外部行为信号,引导模型朝着临床一致的视觉推理方向发展。我们设计了目光引导的推理导航器(GRN),通过双模式探索策略模拟诊断过程,扫描系统性异常定位并深入区域分析。为了平衡专家模仿和自主发现,我们引入了置信度值采样器(CVS),它利用nucleus采样和自适应终止来创建多样但可信的探索路径。最后,双流GRPO优化框架将策略学习和离策略学习信号解耦,缓解奖励同化和熵崩溃。实验表明,MedEyes在多个医学VQA基准测试中实现了平均性能提升+8.5个百分点,验证了MedEyes在构建可信医学AI系统中的潜力。代码可在https://github.com/zhcz328/MedEyes上获得。

英文摘要:

Accurate medical diagnosis often involves progressive visual focusing and iterative reasoning, characteristics commonly observed in clinical workflows. While recent vision-language models demonstrate promising chain-of-thought (CoT) reasoning capabilities via reinforcement learning with verifiable rewards (RLVR), their purely on-policy learning paradigm tends to reinforce superficially coherent but clinically inaccurate reasoning paths. We propose MedEyes, a novel reinforcement learning framework that dynamically models clinician-style diagnostic reasoning by progressively attending to and interpreting relevant medical image regions. By incorporating off-policy expert guidance, MedEyes converts expert visual search trajectories into structured external behavioral signals, guiding the model toward clinically aligned visual reasoning. We design the Gaze-guided Reasoning Navigator (GRN) to emulate the diagnostic process through a dual-mode exploration strategy, scanning for systematic abnormality localization and drilling for detailed regional analysis. To balance expert imitation and autonomous discovery, we introduce the Confidence Value Sampler (CVS), which employs nucleus sampling and adaptive termination to create diverse yet credible exploration paths. Finally, the dual-stream GRPO optimization framework decouples on-policy and off-policy learning signals, mitigating reward assimilation and entropy collapse. Experiments demonstrate that MedEyes achieves an average performance improvement of +8.5pp across multiple medical VQA benchmarks, validating MedEyes's potential in building trustworthy medical AI systems. Code is available at https://github.com/zhcz328/MedEyes.

补充信息

↑