arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

通过OMOP对齐检索教导大型语言模型(LLM)采用ICU医师的临床推理方式可提升跨临床领域的推理能力

Teaching LLMs How ICU Physicians Approach Clinical Reasoning Through OMOP-Aligned Retrieval Improves Reasoning Across Clinical Domains

Miguel Contreras, Scott Siegel, Subhash Nerella, Jessica Sena, Jiaqing Zhang, Heng Sun, Hruday Tej Akkaladevi, Peiyu Lu, Jordan Rosen, Sumit Kapoor, Sasank Desaraju, Grace R. Thompson, Jacob Purcell, Michael Petrauskis, Philip KW. Hong, Meghan Brennan, Sarah Chrabaszcz, Tierra Smith, Ronnie Ren, Michel S. Kabbash, Ceyhun Haziroglu, Rushi Patel, Gabriel Gomez, Charlotte Chaiklin, Randy Leung, Kenneth N. John, Whitman Wiggins, Philip Kayser, Vincent Bird, Maria Bruzzone, Tyler J. Loftus, Azra Bihorac, Parisa Rashidi

arXiv 2608.22622首次发表:更新:

发表机构

University of Florida(佛罗里达大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究开发ICU-REACT推理数据集并微调Clin-REACT模型,使其学习ICU医师的临床推理方式,该模型在五个临床推理基准中表现优于多款模型,可提升跨领域临床推理能力。

AI 中文摘要

临床决策依赖于识别相关患者信息以指导诊断与治疗,这一挑战在数据密集且快速变化的重症监护病房(ICU)中尤为突出。大型语言模型(LLM)可支持该任务,但现有应用与数据集大多侧重表层检索或事实回忆,而非医师用于选择和推理决策相关证据的归纳与演绎推理。我们假设,基于ICU专家推理训练LLM可产生超出重症监护范围的临床推理技能。本文引入ICU-REACT,这是一个由19名医师通过“医师在环”框架开发的推理数据集,用于教导LLM在ICU中执行信息检索与上下文感知的临床推理。利用ICU-REACT,我们对参数规模在80亿至70亿之间、涵盖三个模型系列的Clin-REACT模型进行了微调。在五个临床推理基准测试中,Clin-REACT始终优于其主干模型以及开源通用和医学LLM,性能提升还延伸至脚本一致性测试、下游诊断与治疗等不同任务。这些发现表明,重症监护中的专家推理监督可改善更广泛的临床推理,但在实际临床应用前仍需前瞻性评估。

英文摘要

Clinical decision-making relies on identifying relevant patient information to guide diagnosis and treatment, a challenge that is especially difficult in the data-dense and rapidly changing intensive care unit (ICU). Large language models (LLMs) could support this task. However, existing applications and datasets mostly emphasize surface-level retrieval or factual recall rather than the inductive and deductive reasoning clinicians practice to select and reason over decision-relevant evidence. We hypothesized that training LLMs on expert ICU reasoning could yield clinical reasoning skills that generalize beyond critical care. Here we introduce ICU-REACT, a reasoning dataset developed with 19 clinicians through a clinician-in-the-loop framework to teach LLMs to perform information retrieval and context-aware clinical reasoning in the ICU. Using ICU-REACT, we fine-tuned Clin-REACT models spanning 8B-70B parameters and three model families. Across five clinical reasoning benchmarks, Clin-REACT consistently outperformed its backbone models and open-source general-purpose and medical LLMs. Gains extended to different tasks including script concordance tests, and downstream diagnosis and treatment tasks. These findings suggest that expert reasoning supervision in critical care can improve broader clinical reasoning, although prospective evaluation is needed before real-world clinical use.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑