arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2509.15279cs.LGcs.CL

Fleming-R1:通过强化学习迈向专家级医学推理

Fleming-R1: Toward Expert-Level Medical Reasoning via Reinforcement Learning

  • Ubiquant

机构由 AI 辅助整理,请以论文原文为准。

Chi Liu, Derek Li, Yan Shu, Robin Chen, Derek Duan, Teng Fang, Bryan Dai

更新

AI总结:

Fleming-R1通过面向推理的数据策略、思维链冷启动和两阶段可验证奖励强化学习,在医学推理任务中以参数高效的方式达到接近GPT-4o的性能。

AI中文摘要:

尽管大语言模型在医学应用中展现出潜力,但由于既需要准确的答案又需要透明的推理过程,实现专家级临床推理仍然具有挑战性。为应对这一挑战,我们提出了Fleming-R1,一个通过三项互补创新实现可验证医学推理的模型。首先,我们的面向推理的数据策略(RODS)将精选的医学问答数据集与知识图谱引导的合成相结合,以提高对代表性不足的疾病、药物和多跳推理链的覆盖。其次,我们采用思维链(CoT)冷启动从教师模型中蒸馏高质量的推理轨迹,建立稳健的推理先验。第三,我们实现了一个两阶段的基于可验证奖励的强化学习(RLVR)框架,使用分组相对策略优化,在巩固核心推理技能的同时,通过自适应困难样本挖掘针对持续存在的失败模式。在多个医学基准测试中,Fleming-R1带来了显著的参数效率提升:7B变体超越了规模大得多的基线模型,而32B模型达到了与GPT-4o接近的性能,并持续优于强大的开源替代方案。这些结果表明,结构化的数据设计、面向推理的初始化和可验证的强化学习能够推动临床推理超越简单的准确率优化。我们公开发布Fleming-R1,以促进医学人工智能领域透明、可复现和可审计的进展,从而实现在高风险临床环境中更安全的部署。

英文摘要:

While large language models show promise in medical applications, achieving expert-level clinical reasoning remains challenging due to the need for both accurate answers and transparent reasoning processes. To address this challenge, we introduce Fleming-R1, a model designed for verifiable medical reasoning through three complementary innovations. First, our Reasoning-Oriented Data Strategy (RODS) combines curated medical QA datasets with knowledge-graph-guided synthesis to improve coverage of underrepresented diseases, drugs, and multi-hop reasoning chains. Second, we employ Chain-of-Thought (CoT) cold start to distill high-quality reasoning trajectories from teacher models, establishing robust inference priors. Third, we implement a two-stage Reinforcement Learning from Verifiable Rewards (RLVR) framework using Group Relative Policy Optimization, which consolidates core reasoning skills while targeting persistent failure modes through adaptive hard-sample mining. Across diverse medical benchmarks, Fleming-R1 delivers substantial parameter-efficient improvements: the 7B variant surpasses much larger baselines, while the 32B model achieves near-parity with GPT-4o and consistently outperforms strong open-source alternatives. These results demonstrate that structured data design, reasoning-oriented initialization, and verifiable reinforcement learning can advance clinical reasoning beyond simple accuracy optimization. We release Fleming-R1 publicly to promote transparent, reproducible, and auditable progress in medical AI, enabling safer deployment in high-stakes clinical environments.

↑