发表机构
University of Florida(佛罗里达大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究开发ICU-REACT推理数据集并微调Clin-REACT模型,使其学习ICU医师的临床推理方式,该模型在五个临床推理基准中表现优于多款模型,可提升跨领域临床推理能力。
AI 中文摘要
临床决策依赖于识别相关患者信息以指导诊断与治疗,这一挑战在数据密集且快速变化的重症监护病房(ICU)中尤为突出。大型语言模型(LLM)可支持该任务,但现有应用与数据集大多侧重表层检索或事实回忆,而非医师用于选择和推理决策相关证据的归纳与演绎推理。我们假设,基于ICU专家推理训练LLM可产生超出重症监护范围的临床推理技能。本文引入ICU-REACT,这是一个由19名医师通过“医师在环”框架开发的推理数据集,用于教导LLM在ICU中执行信息检索与上下文感知的临床推理。利用ICU-REACT,我们对参数规模在80亿至70亿之间、涵盖三个模型系列的Clin-REACT模型进行了微调。在五个临床推理基准测试中,Clin-REACT始终优于其主干模型以及开源通用和医学LLM,性能提升还延伸至脚本一致性测试、下游诊断与治疗等不同任务。这些发现表明,重症监护中的专家推理监督可改善更广泛的临床推理,但在实际临床应用前仍需前瞻性评估。
英文摘要
Clinical decision-making relies on identifying relevant patient information to guide diagnosis and treatment, a challenge that is especially difficult in the data-dense and rapidly changing intensive care unit (ICU). Large language models (LLMs) could support this task. However, existing applications and datasets mostly emphasize surface-level retrieval or factual recall rather than the inductive and deductive reasoning clinicians practice to select and reason over decision-relevant evidence. We hypothesized that training LLMs on expert ICU reasoning could yield clinical reasoning skills that generalize beyond critical care. Here we introduce ICU-REACT, a reasoning dataset developed with 19 clinicians through a clinician-in-the-loop framework to teach LLMs to perform information retrieval and context-aware clinical reasoning in the ICU. Using ICU-REACT, we fine-tuned Clin-REACT models spanning 8B-70B parameters and three model families. Across five clinical reasoning benchmarks, Clin-REACT consistently outperformed its backbone models and open-source general-purpose and medical LLMs. Gains extended to different tasks including script concordance tests, and downstream diagnosis and treatment tasks. These findings suggest that expert reasoning supervision in critical care can improve broader clinical reasoning, although prospective evaluation is needed before real-world clinical use.