发表机构
Duke University(杜克大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
EviGen提出三层框架,通过预测性检索、LLM生成和过程监督验证,提升临床推理的可验证性与忠实度,优于现有基线。
AI 中文摘要
纵向电子健康记录(EHRs)捕获了患者多年来的病史,涵盖笔记、编码、实验室检查和程序,并包含推理可能临床结局所需的证据。然而,临床医生全面审阅这些记录是不切实际的,而基于LLM的处理成本高昂且往往不可靠,会遗漏一些相关观察结果,同时产生其他幻觉。因此,我们提出EviGen,一个用于可验证临床推理生成的三层框架,以解决这些挑战。第一层是患者条件检索器,使用可学习查询来查找对临床结局具有预测性(而不仅仅是文本相关性)的证据,并根据预测归因分数对其进行排序。第二层是LLM生成器,它消耗这些排序后的证据作为脚手架,以生成基于检索到的文本片段的临床推理。第三层是过程监督验证器,它在推理步骤级别检查生成的推理,标记不可靠的声明。在三个医学预测数据集上,EviGen在预测性能和推理忠实度上优于全上下文LLM和RAG基线,并且在可用性评估中受到临床审阅者的青睐。
英文摘要
Longitudinal electronic health records (EHRs) capture years of patient history across notes, codes, labs, and procedures, and contain evidence needed to reason about likely clinical outcomes. However, comprehensive clinician review of these records is impractical, and LLM-based processing is costly and often unreliable, missing some relevant observations while hallucinating others. We therefore propose EviGen, a three-layer framework for verifiable clinical rationale generation that addresses these challenges. The first layer is a patient-conditioned retriever that uses learnable queries to find evidence predictive of, not just textually relevant to, a clinical outcome and ranks it by prediction attribution scores. The second layer is an LLM generator that consumes this ranked evidence as a scaffold to produce a clinical rationale grounded in the retrieved spans. The third layer is a process-supervised verifier that checks the generated rationale at the reasoning-step level, flagging unreliable claims. Across three medical prediction datasets, EviGen improves prediction performance and rationale faithfulness over full-context LLM and RAG baselines, and is preferred by clinical reviewers in a usability evaluation.
CommentsAccepted to Findings of EMNLP 2026. 29 pages, 4 figures, 23 tables