MedUPS:利用大语言模型辅助罕见医疗病例的诊断
MedUPS: Towards Diagnostic Assistance in Uncommon Medical Cases with Large Language Models
浏览论文内容
中文总结 AI 辅助
本研究提出MedUPS对齐框架及MedUPSQA数据集,通过强化学习使大语言模型对齐临床中游决策,提升了不同规模模型的下一步临床决策准确率,且小模型表现优于部分大模型。
中文摘要 AI 辅助
罕见及超出诊疗指南的病例对临床决策支持构成挑战,因为医师需在诊断不确定性下作出一系列管理决策,且很少能一次性看到完整病例。大多数面向医学的大语言模型(LLM)基准仅对最终诊断评分,但临床诊疗更多取决于下一步恰当行动:需开具的下一项检查、需进行的影像检查、需邀请的专科医师或需排查的鉴别诊断。我们推出MedUPSQA数据集,该数据集包含从5535份真实病例报告中构建的21874个中游临床决策点,以及MedUPS对齐框架,该框架在患者病程中逐步展开时对模型的中间决策进行监督。我们将自由文本病例陈述分割为按时间顺序排列、逐步累积的临床片段,并利用强化学习(GRPO)结合外部LLM作为评判者的奖励,使模型对齐以预测下一步行动。这一目标反映了临床医师实际接诊患者的方式——从逐步累积的证据向前推理至下一步决策,而非仅确定最终标签。在三种主干模型上,中游对齐使Qwen3.6-27B的下一步准确率从55.2%提升至66.7%,Qwen3.5-9B从47.2%提升至57.8%,HuatuoGPT-3-8B从37.8%提升至44.4%(均含95%置信区间)。在我们测试的多个模型规模中,该目标比模型规模更能提升准确率,较小模型在我们评估中超越了更大的前沿模型。我们还在中游任务上训练了监督微调(SFT)基线,SFT使所有主干模型均优于基础模型,表明目标框架独立于优化器携带有效信号。我们发布该数据集、代码及对齐后的检查点。
英文摘要
Uncommon and off-guideline cases are difficult for clinical decision support, because physicians must make a series of management decisions under diagnostic uncertainty and rarely see the full case at once. Most large language model (LLM) benchmarks for medicine score only the final diagnosis, yet much of clinical care turns on the next appropriate action: the next test to order, the imaging study to obtain, the specialist to involve, or the differential to pursue. We introduce MedUPSQA, a dataset of 21,874 mid-stream clinical decision points built from 5,535 real case reports, and MedUPS, an alignment framework that supervises models on these intermediate decisions as they unfold along a patient's trajectory. We segment free-text case presentations into chronologically ordered, accumulating clinical chunks and align models to predict the next step with reinforcement learning (GRPO), using an external LLM-as-a-Judge reward. This objective mirrors how clinicians actually meet patients, reasoning forward from accumulating evidence toward the next decision, rather than committing to a final label. Across three backbones, mid-stream alignment raises next-step accuracy from 55.2 to 66.7 for Qwen3.6-27B, from 47.2 to 57.8 for Qwen3.5-9B, and from 37.8 to 44.4 for HuatuoGPT-3-8B, with 95% CI. In several model scales we test the objective improves accuracy more than scale, with smaller models surpassing larger, frontier models we evaluate. We further train supervised fine-tuning (SFT) baselines on the mid-stream task, SFT improves all backbones above base, indicating the target framwork carries signal independently of the optimizer. We release the dataset, code, and aligned checkpoints.