Fathom-Vaidya:基于规则奖励推进医学推理
Fathom-Vaidya: Advancing Medical Reasoning with Rubric-Based Rewards
浏览论文内容
中文总结 AI 辅助
提出基于合成数据与规则强化学习的顺序训练框架,分别提升诊断推理与临床推理,在MedXpertQA提升超10%,30B模型在HealthBench-Hard达50.1%,超越GPT-5等基线。
中文摘要 AI 辅助
在医疗保健领域部署大型语言模型(LLM)需要在两个互补维度上具备稳健性能:诊断推理——从临床数据推断患者状况以产生诊断的收敛性、证据驱动任务;以及临床医疗推理——在多轮临床互动中所需的更广泛的导航性判断,用于沟通、规划和适应,其中可能不存在单一正确答案。最近的基准测试(如HealthBench和MedXpertQA)揭示了这两个领域的持续弱点,暴露了复杂诊断场景中的失败以及情境化、以患者为中心对话中的局限性。我们引入了一个顺序训练框架,利用合成数据和基于规则的强化学习针对这些方面。首先,我们使用源自MedBullets的问题,通过规则和规则引导的强化学习(RL)来改进诊断推理。然后,我们转向临床推理,生成5.3k个合成多轮场景,每个场景配有多维规则以全面评估响应。该方法在MedXpertQA上取得了超过10%的提升,我们的30B模型在HealthBench-Hard上达到了50.1%的准确率,超越了包括GPT-5(思考模式)在内的专有基线。我们的结果表明,针对性的合成数据集和基于规则的训练可以系统地改进医疗LLM中的诊断和交互式临床推理。
英文摘要
Deploying Large Language Models (LLMs) in healthcare requires robust performance across two complementary dimensions - diagnostic reasoning: the convergent, evidence-driven task of inferring a patient's condition from clinical data to produce a diagnosis, and clinical healthcare reasoning: the broader, navigational judgment required to communicate, plan, and adapt across multi-turn clinical interactions where a single correct answer may not exist. Recent benchmarks such as HealthBench and MedXpertQA reveal persistent weaknesses in both areas, exposing failures in complex diagnostic scenarios and limitations in contextual, patient-centered dialogue. We introduce a sequential training framework that targets these facets using synthetic data and rubric-based reinforcement learning. First, we improve diagnostic reasoning using MedBullets-derived questions with rule- and rubric-guided Reinforcement Learning (RL). We then shift to clinical reasoning by generating 5.3k synthetic multi-turn scenarios, each paired with multi-dimensional rubrics to comprehensively assess the response. This approach yields over 10% improvement on MedXpertQA, and our 30B model achieves 50.1% accuracy on HealthBench-Hard, surpassing proprietary baselines including GPT-5 (thinking). Our results show that targeted synthetic datasets and rubric-based training can systematically improve both diagnostic and interactive clinical reasoning in medical LLMs.
发表机构
- Fractal AI Research(Fractal AI研究院)
机构由 AI 辅助整理,请以论文原文为准。