AI 中文总结
IyawoBench v2.0评估尼日利亚基层医疗的大型语言模型临床分诊,提出含三类失效模式的框架及新指标,发现前沿模型存缺陷,最优模型随部署场景变化,为中低收入国家临床AI选择提供基准。
AI 中文摘要
大型语言模型正被部署为中低收入国家的临床分诊工具,这些国家训练有素的医生匮乏。然而,现有的安全指标会产生误导性置信度:在二元“未将急诊患者送回家”安全措施上得分为100%的模型,仍可能表现出系统性失效模式,导致其无法大规模部署。我们提出IyawoBench v2.0,这是对大型语言模型临床分诊的扩展诊断评估,基于尼日利亚19家基层医疗中心1200例真实患者 encounters 生成的200个合成 vignette 开展。我们引入一个包含14个定义和2个定理的形式数学框架,将分诊安全分解为三种不同的失效模式:保守升级偏差、系统性降级偏差和中层不稳定。我们提出升级偏差指数和预期部署成本作为新指标,可揭示被传统准确率和灵敏度分数掩盖的失效模式。在三个前沿模型(Claude Sonnet 4.6、Llama 3.3 70B、Llama 3.1 8B)及五个朴素基准上的评估显示:(1)所有三个模型均表现出至少一种形式化失效模式;(2)传统灵敏度指标掩盖了Llama 3.1 8B存在的77个百分点的分诊不足缺口;(3)最优模型在三种部署场景(以急诊为中心、系统可持续性、平衡)中各不相同,表明单排名基准不足以用于中低收入国家临床AI的选择。IyawoBench v2.0提供了一个严格的基准和可迁移至任何分诊式临床AI评估的诊断框架,所有代码、数据和分析流程均公开可用。
英文摘要
Large language models are being deployed as clinical triage tools in low and middle income countries where trained physicians are scarce. Existing safety metrics, however, produce misleading confidence: models scoring 100% on binary "did not send an emergency home" safety measures may nevertheless exhibit systematic failure modes that render them undeployable at scale. We present IyawoBench v2.0, an extended diagnostic evaluation of large language model clinical triage on 200 synthetic vignettes derived from 1,200 real patient encounters at 19 Nigerian primary health centres. We introduce a formal mathematical framework comprising fourteen definitions and two theorems that decompose triage safety into three distinct failure modes: Conservative Escalation Bias, Systematic Downgrade Bias, and Middle-Tier Instability. We propose the Escalation Bias Index and Expected Deployment Cost as novel metrics that expose failure modes hidden by conventional accuracy and sensitivity scores. Evaluated on three frontier models (Claude Sonnet 4.6, Llama 3.3 70B, Llama 3.1 8B) plus five naive baselines, we show that: (1) all three models exhibit at least one formal failure mode; (2) traditional sensitivity metrics conceal a 77 percentage point under-triage gap in Llama 3.1 8B; (3) the optimal model varies across three deployment scenarios (Emergency-Focused, System-Sustainability, Balanced), demonstrating that single-ranking benchmarks are inadequate for LMIC clinical AI selection. IyawoBench v2.0 provides both a rigorous benchmark and a diagnostic framework transferable to any triage-style clinical AI evaluation. All code, data, and analysis pipelines are publicly available.