发表机构
Duke University; Columbia University Medical Center; Vanderbilt University Medical Center(杜克大学; 哥伦比亚大学医学中心; 范德堡大学医学中心)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究旨在检测电子健康记录中文档不一致性,采用两阶段LLM管道对3000份出院小结进行分析,发现多种类型不一致及反复出现的失败模式,建立了方法基础和概念框架以指导后续大规模分析。
AI 中文摘要
目的:表征通用领域大语言模型(LLM)能从真实世界出院小结中发现的内部文档不一致类型,并识别限制大规模可靠性的反复出现的失败模式。材料和方法:我们将两阶段LLM管道——开放式候选识别(Gemini 2.5 Pro),随后是上下文验证(Gemini 2.5 Flash)——应用于3000份随机抽样的MIMIC-IV-Note出院小结。管道输出的一个子集随后由临床专家手动审查。结果:我们的管道发现了3460个候选不一致,影响69.7%的入院病例。代表性例子涵盖人口统计学、过敏、程序、诊断、实验室、药物和护理计划领域,对临床推理或患者安全有直接影响。专家审查还揭示了在验证需要时间推理、不断演变的诊断背景或模型本身不具备的门诊处方惯例知识时出现的反复出现的失败模式。讨论:检测高度依赖上下文:许多标记的对需要将每个陈述锚定到其源部分和临床领域,然后评估冲突是否反映真正的矛盾或缺失的上下文。我们提出了一个跨越严格矛盾和模糊性的分级本体,其模式按类别、部分、领域和不一致轴对每个标记案例进行表征。结论:这项形成性研究建立了一个方法基础和概念框架,以指导后续经过验证的大规模电子健康记录不一致性分析。
英文摘要
Objective: To characterize the kinds of internal documentation inconsistencies a general-domain large language model (LLM) can surface from real-world discharge summaries, and to identify recurring failure modes that limit reliability at scale. Materials and Methods: We applied a two-stage LLM pipeline---open-ended candidate identification (Gemini 2.5 Pro) followed by context-grounded verification (Gemini 2.5 Flash)---to 3,000 randomly sampled MIMIC-IV-Note discharge summaries. A subset of the pipeline output was then reviewed manually by clinical experts. Results: Our pipeline surfaced 3,460 candidate inconsistencies, affecting 69.7% of admissions. Representative examples spanned demographics, allergies, procedures, diagnoses, laboratory, medications, and care-planning domains, with direct implications for clinical reasoning or patient safety. Expert review also revealed recurring failure modes that arise when verification requires temporal reasoning, evolving-diagnosis context, or knowledge of outpatient-prescribing conventions the model does not natively possess. Discussion: Detection is highly context-dependent: many flagged pairs require anchoring each statement to its source section and clinical domain, then assessing whether the conflict reflects a true contradiction or missing context. We propose a graded ontology spanning strict contradiction and ambiguity, with a schema characterizing each flagged case by category, section, domain, and inconsistency axis. Conclusion: This formative study establishes a methodological foundation and conceptual framework to guide subsequent validated, large-scale EHR-inconsistency analysis.