印度母婴护理的可审计紧急分诊
Auditable Emergency Triage for Maternal and Newborn Care in India
浏览论文内容
中文总结 AI 辅助
针对母婴护理紧急分诊,提出LLM提取症状与规则引擎结合的可审计系统,提升召回率与F1,支持临床专家独立添加规则。
中文摘要 AI 辅助
在诺拉健康(Noora Health),我们的护士每月通过基于WhatsApp的服务处理超过50,000条医疗查询,该服务为照护者提供按需支持。她们最关键的任务是紧急分诊:决定哪些查询需要立即的面对面关注。为了支持她们,我们构建了一个系统,使用大型语言模型(LLM)来分类一条消息是否为紧急情况,并提供理由以增强可解释性。但该系统是不透明的:分析错误意味着阅读每条消息的推理链,这在我们的规模下是不可行的。提示更改意味着重新运行完整的评估以防止回归,这既昂贵又在操作上具有挑战性。临床医生遵循决策树来做出这一判断,但该决策树从未被记录或传递给模型,模型依赖于一组扁平的 danger signs 列表。为了解决这些问题,我们将分诊分解为两个步骤:LLM使用临床医生编写的词汇表从查询中提取规范症状和患者背景,然后一个确定性规则引擎捕获指示紧急情况的场景。我们表明,新系统将召回率从0.565提高到0.810,F1分数从0.606提高到0.702,其中结构化规则驱动了大部分准确性提升,而分解提供了可审计性:临床专家可以检查新系统的每个阶段,以查看查询是否被错误翻译、症状是否被错误提取、患者背景是否被错误推断,或必要的规则是否缺失。他们可以独立添加新规则而不会导致回归,并避免运行昂贵的评估。自部署以来,新系统已对152,421条患者查询进行了分诊,并标记了28,535条(18.7%)为紧急情况。过度升级率为17.8%,且未增加漏诊紧急情况。临床医生自部署以来还添加了48条新规则,这证明了我们着手构建的更快的纠正循环。
英文摘要
At Noora Health, our nurses answer more than 50,000 medical queries per month on our WhatsApp-based service that provides caregivers with on-demand support. Their most time-critical task is emergency triage: deciding which queries need immediate in-person attention. To support them, we built a system that uses a large language model (LLM) to classify whether a message is an emergency and provide a rationale for interpretability. But the system was opaque: analyzing mistakes meant reading reasoning chains for each message, which is infeasible at our scale. Prompt changes meant re-running a full evaluation to prevent regressions, which was both costly and operationally challenging. Clinicians follow a decision tree to make this call, but it was never documented or passed to the model, which relied on a flat list of danger signs. To address these issues, we decomposed triage into two steps: an LLM extracts canonical symptoms and patient context from the query using a clinician-authored vocabulary, and a deterministic rule engine captures the scenarios that indicate an emergency. We show that the new system raised recall from 0.565 to 0.810 and F1 from 0.606 to 0.702, with structured rules driving most of the accuracy gains while the decomposition provides auditability: clinical experts can inspect each stage of the new system to see whether the query was mistranslated, symptoms were incorrectly extracted, patient context was wrongly inferred, or the necessary rules were missing. They can add new rules independently without causing regressions and avoid running costly evaluations. Since deployment, the new system has triaged 152,421 patient queries and flagged 28,535 (18.7%) as emergencies. The over-escalation rate has been 17.8%, without any increase in missed emergencies. Clinicians have also added 48 new rules since deployment, evidence of the faster correction loop we set out to build.
发表机构
- Noora Health
- The Agency Fund
机构由 AI 辅助整理,请以论文原文为准。