发表机构
Google DeepMind; Google Research; Houston Methodist Hospital; Trinity Health Group; Stanford Oncology Partners; St. Luke Hospital(谷歌DeepMind; 谷歌研究院; 休斯顿卫理公会医院; 三一健康集团; 斯坦福肿瘤学伙伴; 圣卢克医院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究提出 ResidencyRL,通过多轮强化学习训练临床 AI 智能体,在模拟临床环境中提升诊断准确性、降低漏报率,且能力可迁移至多个医学基准测试,为临床 AI 发展提供了新路径。
AI 中文摘要
在医学教育中,医师通过住院医师培训将学术知识转化为临床专长,该过程包含数千次诊疗 encounter,涵盖各类反馈来源及逐步提升的自主性。临床推理很大程度依赖于医患对话 encounter,临床医师需从中采集病史、完善诊断假设,并在不确定性下制定诊疗方案。尽管大语言模型(LLM)在静态医学基准测试中表现出色,但优化完整临床决策序列的方法仍未充分发展。我们提出 ResidencyRL,这是一种强化学习(RL)方法,用于通过模拟多轮临床 encounter 训练临床人工智能(AI)智能体,每个轨迹最多包含 60 轮对话和 8 次工具调用。ResidencyRL 将策略智能体与能够产生复杂对抗性行为的 LLM 模拟器配对,依据与诊断准确性、诊疗质量、沟通、文档记录及安全性对齐的结构化奖励进行训练。在保留的评估中,ResidencyRL 智能体在对抗条件下将诊断准确性提高了 7.0%(88.0% 对比 81.0%),并将漏报危险信号的比率降低了 31%,证明其能严格缓解过早闭合问题。盲法专家临床医师验证了这些改进,在 87.6% 的并排比较中更倾向于该训练后的智能体。操作能力可迁移至未见的基准测试:该智能体在 AMIE 多访视基准测试的全部六个临床维度上均优于基础模型,并在 AgentClinic 和 CRAFT-MD 上表现出一致的方向性改进。我们的研究结果表明,通过模拟中的多轮 RL 可有效学习序列临床决策,从而产生稳健、可泛化的能力,为实现临床精通铺平道路。仍需对真实世界工作流程进行前瞻性验证以确立临床实用性。
英文摘要
In medical education, physicians convert academic knowledge into clinical expertise through residency: years of training across thousands of encounters, with diverse sources of feedback and progressively greater autonomy. Much of clinical reasoning relies on the patient encounter, a dialogue in which a clinician elicits history, refines diagnostic hypotheses, and decides management under uncertainty. While large language models (LLMs) excel on static medical benchmarks, methods to optimize the full sequence of clinical decisions remain underdeveloped. We present ResidencyRL, a reinforcement learning (RL) method for training clinical artificial intelligence (AI) agents through simulated multi-turn clinical encounters (up to 60 dialogue turns and 8 tool calls per trajectory). ResidencyRL pairs the policy agent with LLM simulators capable of complex, adversarial behaviors, training against a structured reward aligned to diagnostic accuracy, management quality, communication, documentation, and safety. On held-out evaluations, the ResidencyRL agent improves diagnostic accuracy by 7.0% under adversarial conditions (88.0% vs. 81.0%) and reduces missed red flag rates by 31%, demonstrating rigorous mitigation of premature closure. Blinded expert clinicians validated these gains, preferring the trained agent in 87.6% of side-by-side comparisons. The procedural competencies transfer to unseen benchmarks: the agent outperforms the base model across all six clinical axes of the AMIE multi-visit benchmark, and shows consistent directional improvements on AgentClinic and CRAFT-MD. Our findings demonstrate that sequential clinical decision-making can be effectively learned through multi-turn RL in simulation, yielding robust, generalizable capabilities, paving the way towards clinical mastery. Prospective validation with real-world workflows remains necessary to establish clinical utility.