急诊科再就诊质量审查筛查:探索人类决策与人工智能支持
Emergency Department Revisit Quality Review Screening: Exploring Human Decision-Making and Artificial Intelligence Support
浏览论文内容
中文总结 AI 辅助
本研究探索急诊科再就诊质量审查中人类决策与GPT-4等AI支持,提出基于知识图谱的KGA算法,实现高阳性预测值筛查,有望扩大筛查范围并减少工作量。
中文摘要 AI 辅助
背景:急诊科(ED)再就诊通常会被纳入质量保证审查,但此类审查往往受限(例如,仅审查48-72小时内的再就诊),以提高可操作发现的产出率,同时尽量减少病历审查负担。这些限制可能导致错失质量改进机会。方法:我们开展了一项探索性、回顾性研究,随机选取了一家多医院医疗系统内,在初次就诊后1-14天内再次就诊于同一医疗系统的急诊科病例。仅根据每次就诊的主要诊断,评估者(2-3名临床医生和GPT-4大语言模型[LLM])评估了诊断对的特征,包括“目标”:即该诊断对是否需要进一步评估。基于评估者反应分析,我们创建了一种利用LLM填充的知识图谱(“KGA”)的算法,用于自动筛查潜在令人担忧的诊断对,并进行了初步评估。结果:共纳入99个诊断对。GPT-4的响应与临床医生评估者的相关性较差,其将几乎所有(94%)诊断对评为需要随访(比临床医生多4.4-13.3倍)。然而,提示工程使用极少。在临床医生评估者中,再就诊的医学严重程度始终与目标显著相关,而鉴别诊断/并发症复合指标在未调整分析中显著相关,但在调整分析中不显著(尽管统计功效较低)。KGA在至少一名临床医生评估者认为基于诊断对需要进一步评估的情况下,实现了83-100%的阳性预测值。结论:这些结果可为利用ChatGPT等LLM改进筛查的后续步骤提供参考。需要进一步研究来验证这项初步工作的发现,即KGA可能在不显著增加审查者工作量的情况下,扩大筛查范围和产出率。
英文摘要
Background: Emergency Department (ED) return visits are commonly reviewed for quality assurance, but are often limited (e.g., to revisits within 48-72 hours) to increase actionable finding yield while minimizing chart review burden. Those limitations may lead to missed quality improvement opportunities. Methods: We conducted an exploratory, retrospective study of randomly selected ED visits to a multihospital health system having an ED revisit within 1-14 days to the same health system. Given only each visit's primary diagnosis, raters (2-3 clinicians and GPT-4 large language model [LLM]) assessed characteristics of the diagnosis pairs, including the "target": whether a pair warranted further assessment. Informed by rater response analyses, an algorithm leveraging an LLM-populated knowledge graph ("KGA") was created to automatically screen for potentially concerning pairs, then preliminarily assessed. Results: 99 diagnosis pairs were included. GPT-4 responses poorly correlated to clinician raters, rating nearly all (94%) pairs as warranting follow-up (4.4-13.3 times more than clinicians). However, prompt engineering was minimal. Among clinician raters, revisit medical gravity was consistently significantly associated with the target, while a differential diagnosis/complication composite was significantly associated on unadjusted, but not adjusted (though less powered) analysis. The KGA achieved 83-100% positive predictive value for at least one clinician rater determining further assessment was warranted based on the diagnosis pair. Conclusion: These results can inform next steps for improving screening with LLMs like ChatGPT. Further research is warranted to validate this preliminary work's finding that the KGA may enable enhancing the scope and yield of screening without substantially increasing reviewer workload.
发表机构
- OSF HealthCare(OSF医疗保健)
- University of Illinois College of Medicine at Peoria(伊利诺伊大学皮奥里亚医学院)
- Saint Anthony College of Nursing(圣安东尼护理学院)
- Thomas Jefferson University's Sidney Kimmel College of Medicine(托马斯·杰斐逊大学西德尼·金梅尔医学院)
- Keylog Solutions LLC(Keylog解决方案有限责任公司)
机构由 AI 辅助整理,请以论文原文为准。