发表机构
Virginia Tech; University of Washington(弗吉尼亚理工大学; 华盛顿大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究对十个开源大语言模型进行反事实审计,评估其在儿科急诊分诊中的偏见,发现微调模型敏感性最低,支持反事实审计作为临床部署前的公平性评估框架。
AI 中文摘要
急诊科(ED)分诊是一项高风险优先排序任务,其中人口统计学、社会经济和系统背景信息可能不当影响病情严重程度分级。尽管开源大语言模型(LLMs)日益被视为用于本地化、隐私保护的临床决策支持工具,但反事实偏见在不同模型家族、规模、医学领域模型及领域适应模型中的变化情况仍不清楚。我们针对十个开源大语言模型在儿科急诊严重程度指数(ESI)预测中进行了比较性反事实审计。基于真实及手册式临床案例,我们构建了成对的反事实变体,每次仅改变一个注入的人口统计学、社会经济、医疗获取、行为、社会或系统背景变量,同时保持临床表现不变。模型包括Qwen2.5-7B、Qwen2.5-14B-Instruct、经QLoRA微调的Qwen2.5-7B、MedGemma系列、MedLLaMA2-7B、GPT-OSS-20B和GPT-OSS-120B。我们测量了任何反事实偏移、分诊不足、分诊过度、超过一个ESI级别的偏移、平均偏移和平均绝对偏移。反事实敏感性差异显著,并未随模型规模增大或医学领域预训练而一致降低。微调的Qwen2.5-7B显示出最低的整体敏感性,其任何偏移率为5.27%,平均绝对偏移为0.0534,而基础模型分别为16.02%和0.1706。几个更大或医学领域的模型显示出更显著的偏移。分层和相关性分析进一步揭示了临床上重要的方向性和聚合率所隐藏的共同失败模式。这些发现支持反事实审计作为一种轻量级、临床可解释的框架,用于在临床部署前比较开源大语言模型中的公平性风险。
英文摘要
Emergency department (ED) triage is a high-stakes prioritization task in which demographic, socioeconomic, and system-context information may improperly influence acuity assignment. Although open-source large language models (LLMs) are increasingly considered for local and privacy-preserving clinical decision support, it remains unclear how counterfactual bias varies across model families, sizes, medical-domain models, and domain-adapted models. We present a comparative counterfactual audit of ten open-source LLMs for pediatric Emergency Severity Index (ESI) prediction. Starting from real and handbook-style clinical vignettes, we construct paired counterfactual variants that change only one injected demographic, socioeconomic, healthcare-access, behavioral, social, or system-context variable while holding the clinical presentation fixed. Models include Qwen2.5-7B, Qwen2.5-14B-Instruct, a QLoRA fine-tuned Qwen2.5-7B, MedGemma variants, MedLLaMA2-7B, GPT-OSS-20B, and GPT-OSS-120B. We measure any counterfactual shift, undertriage, overtriage, shifts greater than one ESI level, mean shift, and mean absolute shift. Counterfactual sensitivity varied substantially and did not consistently decrease with larger model size or medical-domain pretraining. The fine-tuned Qwen2.5-7B showed the lowest overall sensitivity, with a 5.27% any-shift rate and mean absolute shift of 0.0534, versus 16.02% and 0.1706 for the base model. Several larger or medical-domain models showed more significant shifts. Stratified and correlation analyses further revealed clinically important directionality and shared failure patterns hidden by aggregate rates. These findings support counterfactual auditing as a lightweight, clinically interpretable framework for comparing fairness risks in open-source LLMs before clinical deployment.