MamaBench:通过反事实临床扰动对母婴健康诊断中的大语言模型稳健性进行基准测试
MamaBench: Benchmarking LLM Robustness in Maternal and Child Health Diagnosis through Counterfactual Clinical Perturbation
浏览论文内容
中文总结 AI 辅助
研究针对大语言模型在母婴健康诊断中缺乏对临床相似表现区分能力的问题,提出MamaBench基准及EA - RAG方法,通过实验揭示基础准确率高估稳健准确率现象,EA - RAG有效降低BTR,提升稳健准确率,为临床人工智能反事实稳健性研究提供参考。
中文摘要 AI 辅助
大语言模型在医学基准测试中取得高分,但这些基准测试孤立地评估每个问题,无法衡量系统是否能区分需要不同干预的临床相似表现。我们引入了MamaBench,首个母婴人工智能反事实基准:涵盖371种病症的217对共434篇专家撰写的临床叙述,通过偏差陷阱率(BTR)评估。我们提出证据锚定检索增强生成(EA - RAG),一种三阶段检索方法。实验表明,在四个前沿大语言模型的八种配置中,基础准确率高估了稳健准确率。EA - RAG在Claude Sonnet 4.6上实现了20.3%的BTR和65.0%的稳健准确率,BTR降低了5.5个百分点且不降低基础准确率。剩余20%的BTR表明临床人工智能中的反事实稳健性仍是一个开放挑战。
英文摘要
Large language models achieve strong scores on medical benchmarks, yet these benchmarks evaluate each question in isolation, providing no measure of whether a system can distinguish clinically similar presentations requiring different interventions. We introduce MamaBench, the first counterfactual benchmark for maternal and paediatric AI: 434 expert-authored clinical narratives in 217 pairs across 371 pathologies, evaluated via the Bias Trap Rate (BTR), the conditional probability that a model fails the counterfactual given success on the base case. We propose Evidence-Anchored RAG (EA-RAG), a three-stage retrieval method that replaces aggregate similarity with an evidence coverage objective through clinical parameter extraction, coverage auditing, and contrastive sub-queries. Across eight configurations of four frontier LLMs, base accuracy overstates robust accuracy by 16-28 percentage points in every model. EA-RAG achieves 20.3% BTR and 65.0% robust accuracy on Claude Sonnet 4.6, a 5.5 percentage point BTR reduction without degrading base accuracy. The residual 20% BTR confirms that counterfactual robustness in clinical AI remains an open challenge. Keywords: counterfactual evaluation, clinical AI, maternal healthcare, retrieval-augmented generation, diagnostic robustness
发表机构
- Helpmum Africa(非洲帮助妈妈组织)
- University of Ibadan(伊巴丹大学)
机构由 AI 辅助整理,请以论文原文为准。