大规模大语言模型用于罕见病鉴别诊断:从腹部放线菌病到威尔逊病
Rare Disease Differential Diagnosis with Large Language Models at Scale: From Abdominal Actinomycosis to Wilson's Disease
- Curai Health(库赖医疗)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
针对大语言模型罕见病诊断能力不足的问题,提出RareScale框架结合专家系统与LLM,通过生成模拟数据训练候选预测模型辅助黑盒LLM,显著提升575种以上罕见病的Top-5诊断准确率。
AI中文摘要:
大语言模型(LLMs)已在疾病诊断方面展现出令人瞩目的能力。然而,它们在识别罕见病方面的有效性仍是一个悬而未决的问题,而罕见病的诊断本身就更具挑战性。随着LLMs在医疗场景中的应用日益广泛,其在罕见病诊断上的表现至关重要。当初级保健医生仅能通过与患者的对话做出罕见病预后判断、以便采取恰当的后续措施时,这一点尤为关键。为此,已有多款临床决策支持系统被设计用于辅助医疗人员识别罕见病,但这类系统因对常见疾病知识储备不足且使用难度高,实用性受到限制。\n本文提出RareScale框架,将LLMs的知识与专家系统相结合。我们联合使用专家系统和LLM模拟罕见病问诊对话,用这些数据训练一个罕见病候选预测模型。该小型模型输出的候选诊断会作为额外输入提供给黑盒LLM,由其完成最终的鉴别诊断。由此,RareScale实现了罕见病与常见病诊断之间的平衡。我们在从腹部放线菌病到威尔逊病的575种以上罕见病上开展了实验。结果显示,该方法将黑盒LLMs的基线Top-5准确率显著提升了17%以上。我们还发现其候选生成性能优异,例如在gpt-4o生成的对话上准确率达88.8%。
英文摘要:
Large language models (LLMs) have demonstrated impressive capabilities in disease diagnosis. However, their effectiveness in identifying rarer diseases, which are inherently more challenging to diagnose, remains an open question. Rare disease performance is critical with the increasing use of LLMs in healthcare settings. This is especially true if a primary care physician needs to make a rarer prognosis from only a patient conversation so that they can take the appropriate next step. To that end, several clinical decision support systems are designed to support providers in rare disease identification. Yet their utility is limited due to their lack of knowledge of common disorders and difficulty of use. In this paper, we propose RareScale to combine the knowledge LLMs with expert systems. We use jointly use an expert system and LLM to simulate rare disease chats. This data is used to train a rare disease candidate predictor model. Candidates from this smaller model are then used as additional inputs to black-box LLM to make the final differential diagnosis. Thus, RareScale allows for a balance between rare and common diagnoses. We present results on over 575 rare diseases, beginning with Abdominal Actinomycosis and ending with Wilson's Disease. Our approach significantly improves the baseline performance of black-box LLMs by over 17% in Top-5 accuracy. We also find that our candidate generation performance is high (e.g. 88.8% on gpt-4o generated chats).