发表机构
The Institute of Product Leadership; Amazon(产品领导力研究院; 亚马逊)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究通过反事实有效性和人口统计稳健性测试评估六个LLM在MedQA上的表现,发现准确率不足以反映临床可靠性,并揭示了模型间的脆弱性和偏见差异。
AI 中文摘要
临床大语言模型(LLM)评估通常强调答案准确性;然而,仅凭准确性并不能检验反事实一致性或人口统计稳健性。我们使用两项自动化扰动测试,对六个大语言模型在150道MedQA USMLE题目上的表现进行了评估。反事实有效性(CFV)测试要求每个模型做出一个最小且合理的临床改变,使另一个答案成为正确答案。人口统计稳健性测试则在相同的临床情境中添加六种人口统计前缀,并与无人口统计基线进行比较答案和解释。在900次CFV尝试中,228次(25.3%)有效,672次无效。在5400次人口统计比较中,有1097个答案发生改变(20.3%)。自动化评判识别出3128个刻板印象证据标记,其中1932个属于广泛的“其他”类别。MedGemma 27B达到了最高的准确率(87.1%)和CFV(63.3%),最低的答案改变率(16.0%),以及较低的平均解释人口统计不和谐(EDD)分数(0.169)。然而,其准确率仍超过CFV,表明正确答案并不能保证在反事实有效性任务上表现可靠。OpenBioLLM的答案改变率和EDD最高,而GLM的刻板印象标记率最高。这些发现表明,准确率、CFV、答案稳定性、EDD和刻板印象证据捕捉了不同的评估方面。由于所有评判均为自动化,且没有临床医生验证,结果支持安全筛选,但不能确立模型的临床可部署性。
英文摘要
Clinical LLM evaluation often emphasizes answer accuracy; however, accuracy alone does not test counterfactual consistency or demographic robustness. We evaluated six LLMs on 150 MedQA USMLE questions using two automated perturbation tests to assess their performance. The counterfactual validity (CFV) test asked each model to make a minimal, plausible clinical change that would make a different answer correct. The demographic robustness test added six demographic prefixes to the same vignette and compared the answers and explanations with a no demographic baseline. Of the 900 CFV attempts, 228 (25.3 %) were valid and 672 were invalid. Across 5,400 demographic comparisons, 1,097 answers were changed (20.3%). Automated judging identified 3,128 stereotype evidence flags, including 1,932 in the broad Other category. MedGemma 27B achieved the highest accuracy (87.1%) and CFV (63.3%), lowest answer change rate (16.0%), and low mean Explanation Demographic Dissonance (EDD) score (0.169). However, its accuracy still exceeded its CFV, indicating that correct answers do not guarantee reliable performance on the counterfactual validity task. OpenBioLLM had the highest answer change rate and EDD, whereas GLM had the highest stereotype flag rate. These findings show that accuracy, CFV, answer stability, EDD, and stereotype evidence capture different evaluation aspects. Because all judgments were automated and no clinician validation was available, the results support safety screening but do not establish clinical deployability of the model.
CommentsAccepted into AIKP 2026 Conference