通过知识增强的数据合成 eliciting 医学推理:一种半监督强化学习方法
Eliciting Medical Reasoning with Knowledge-enhanced Data Synthesis: A Semi-Supervised Reinforcement Learning Approach
- College of Computer Science and Artificial Intelligence, Fudan University(复旦大学计算机科学与技术学院)
- Shanghai AI Laboratory(上海人工智能实验室)
- CMIC, Shanghai Jiao Tong University(上海交通大学CMIC)
- School of Artificial Intelligence, Shanghai Jiao Tong University(上海交通大学人工智能学院)
- Department of Radiology, Shanghai Sixth People’s Hospital Affiliated to Shanghai Jiao Tong University School of Medicine(上海交通大学医学院附属第六人民医院放射科)
- Institute of Artificial Intelligence for Medicine, Shanghai Jiao Tong University School of Medicine(上海交通大学医学院人工智能医学研究所)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本文提出MedSSR框架,利用稀有疾病知识生成可控推理问题,并通过自监督和监督强化学习提升医学推理能力,在十个医疗基准测试中取得显著提升。
AI中文摘要:
尽管大型语言模型在复杂医疗应用中具有潜力,但其发展受到高质量推理数据稀缺的阻碍。为解决此问题,现有方法通常通过监督微调从大型专有模型中提取链式推理轨迹,然后进行强化学习(RL)。这些方法在罕见疾病等代表性不足的领域改进有限,且生成复杂推理链的成本高昂。为高效提升医学推理能力,我们提出MedSSR,一种医疗知识增强的数据合成和半监督强化学习框架。我们的框架首先利用稀有疾病知识生成可控的推理问题。然后利用策略模型本身生成高质量的伪标签。这使我们能够采用两阶段、内在到外在的训练范式:在伪标签合成数据上进行自监督RL,随后在人工标注的真实数据上进行监督RL。MedSSR能够高效扩展模型训练,无需依赖昂贵的轨迹蒸馏。在Qwen和Llama上的大量实验表明,我们的方法在十个医疗基准测试中优于现有方法,实现了在罕见疾病任务上的最大+5.93%的提升。我们的代码可在https://github.com/tdlhl/MedSSR上获得。
英文摘要:
While large language models hold promise for complex medical applications, their development is hindered by the scarcity of high-quality reasoning data. To address this issue, existing approaches typically distill chain-of-thought reasoning traces from large proprietary models via supervised fine-tuning, then conduct reinforcement learning (RL). These methods exhibit limited improvement on underrepresented domains like rare diseases while incurring substantial costs from generating complex reasoning chains. To efficiently enhance medical reasoning, we propose MedSSR, a Medical Knowledge-enhanced data Synthesis and Semi-supervised Reinforcement learning framework. Our framework first employs rare disease knowledge to synthesize distribution-controllable reasoning questions. We then utilize the policy model itself to generate high-quality pseudo-labels. This enables a two-stage, intrinsic-to-extrinsic training paradigm: self-supervised RL on the pseudo-labeled synthetic data, followed by supervised RL on the human-annotated real data. MedSSR scales model training efficiently without relying on costly trace distillation. Extensive experiments on Qwen and Llama demonstrate that our method outperforms existing methods across ten medical benchmarks, achieving up to +5.93% gain on rare-disease tasks. Our code is available at https://github.com/tdlhl/MedSSR.