arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2506.21615cs.CLcs.AIcs.IR

通过生成增强检索和临床实践指南提升医学诊断

Refine Medical Diagnosis Using Generation Augmented Retrieval and Clinical Practice Guidelines

  • Shandong Normal University(山东师范大学)
  • Flinders University(弗林德斯大学)

机构由 AI 辅助整理,请以论文原文为准。

Wenhao Li, Hongkuan Zhang, Hongwei Zhang, Zhengxu Li, Zengjie Dong, Yafan Chen, Niranjan Bidargaddi, Hong Liu

更新

AI总结:

本研究提出GARMLE-G框架,通过生成增强检索和临床指南整合,提升医学诊断的准确性和临床实用性。

AI中文摘要:

当前的医学语言模型通常从大型语言模型(LLMs)中适应,通常从电子健康记录(EHRs)预测基于ICD代码的诊断,因为这些标签易于获取。然而,ICD代码无法捕捉临床医生用于诊断的细致、富含上下文的推理。临床医生综合多样化的患者数据并参考临床实践指南(CPGs)以做出循证决策。这种不匹配限制了现有模型的临床实用性。我们介绍了GARMLE-G,一种生成增强检索框架,使医学语言模型的输出扎根于权威的CPGs中。与传统检索增强生成方法不同,GARMLE-G通过直接检索权威指南内容而不依赖模型生成文本,实现了无幻觉输出。它(1)整合LLM预测与EHR数据以创建语义丰富的查询,(2)通过嵌入相似性检索相关CPGs知识片段,(3)将指南内容与模型输出融合以生成临床对齐的推荐。开发了一个高血压诊断原型系统,并在多个指标上进行评估,结果显示其检索精度、语义相关性和临床指南一致性优于基于RAG的基线模型,同时保持了轻量级架构,适用于本地化医疗部署。本工作提供了一种可扩展、低成本且无幻觉的方法,将医学语言模型扎根于循证临床实践中,具有广泛临床部署的潜力。

英文摘要:

Current medical language models, adapted from large language models (LLMs), typically predict ICD code-based diagnosis from electronic health records (EHRs) because these labels are readily available. However, ICD codes do not capture the nuanced, context-rich reasoning clinicians use for diagnosis. Clinicians synthesize diverse patient data and reference clinical practice guidelines (CPGs) to make evidence-based decisions. This misalignment limits the clinical utility of existing models. We introduce GARMLE-G, a Generation-Augmented Retrieval framework that grounds medical language model outputs in authoritative CPGs. Unlike conventional Retrieval-Augmented Generation based approaches, GARMLE-G enables hallucination-free outputs by directly retrieving authoritative guideline content without relying on model-generated text. It (1) integrates LLM predictions with EHR data to create semantically rich queries, (2) retrieves relevant CPG knowledge snippets via embedding similarity, and (3) fuses guideline content with model output to generate clinically aligned recommendations. A prototype system for hypertension diagnosis was developed and evaluated on multiple metrics, demonstrating superior retrieval precision, semantic relevance, and clinical guideline adherence compared to RAG-based baselines, while maintaining a lightweight architecture suitable for localized healthcare deployment. This work provides a scalable, low-cost, and hallucination-free method for grounding medical language models in evidence-based clinical practice, with strong potential for broader clinical deployment.

补充信息

↑