卓越还是真实表现?通过动态评估重新思考医疗诊断基准
Inflated Excellence or True Performance? Rethinking Medical Diagnostic Benchmarks with Dynamic Evaluation
- GSAI, Renmin University of China(清华大学人工智能研究院,中国人民大学)
- Tencent Jarvis Lab(腾讯 Jarvis实验室)
- Beijing Key Laboratory of Research on Large Models and Intelligent Governance(北京大型模型与智能治理研究关键实验室)
- Engineering Research Center of Next-Generation Intelligent Search and Recommendation, MOE(下一代智能搜索与推荐工程技术研究中心,教育部)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本文提出DyReMe动态基准,通过生成临床相关干扰因素的咨询式案例,评估医疗诊断模型的鲁棒性,揭示现有模型在临床干扰下的不足。
AI中文摘要:
医疗诊断是一个高风险且复杂的领域,对患者护理至关重要。然而,当前大型语言模型(LLM)的评估仍难以捕捉临床诊断场景的关键挑战。大多数评估依赖于公开考试衍生的基准,导致性能膨胀的污染偏差,并忽视了真实咨询中超出教科书案例的混淆性质。最近的动态评估提供了有希望的替代方案,但通常不足以用于诊断导向的基准测试,对临床相关的混淆因素和可信度的覆盖有限。为了解决这些差距,我们提出了DyReMe,一个医疗诊断的动态基准,提供受控且可扩展的诊断鲁棒性压力测试。与静态考试式问题不同,DyReMe生成新鲜的咨询式案例,包含临床相关的混淆因素,如鉴别诊断和常见误诊因素。它还通过变化表达风格来捕捉异质性的患者描述。除了准确性外,DyReMe还评估LLM在三个额外的临床相关维度:真实性、有用性和一致性。我们的实验表明,这种动态方法产生了更具挑战性的评估,并揭示了在临床混淆诊断设置下现有最先进的LLM的显著不足。这些发现突显了需要更好地评估基于临床相关混淆因素的可信医疗诊断的评估框架的紧迫性。
英文摘要:
Medical diagnostics is a high-stakes and complex domain that is critical to patient care. However, current evaluations of large language models (LLMs) remain limited in capturing key challenges of clinical diagnostic scenarios. Most rely on benchmarks derived from public exams, raising contamination bias that can inflate performance, and they overlook the confounded nature of real consultations beyond textbook cases. Recent dynamic evaluations offer a promising alternative, but often remain insufficient for diagnosis-oriented benchmarking, with limited coverage of clinically grounded confounders and trustworthiness beyond accuracy. To address these gaps, we propose DyReMe, a dynamic benchmark for medical diagnostics that provides a controlled and scalable stress test of diagnostic robustness. Unlike static exam-style questions, DyReMe generates fresh, consultation-style cases that incorporate clinically grounded confounders, such as differential diagnoses and common misdiagnosis factors. It also varies expression styles to capture heterogeneous patient-style descriptions. Beyond accuracy, DyReMe evaluates LLMs on three additional clinically relevant dimensions: veracity, helpfulness, and consistency. Our experiments show that this dynamic approach yields more challenging assessments and exposes substantial weaknesses of stateof-the-art LLMs under clinically confounded diagnostic settings. These findings highlight the urgent need for evaluation frameworks that better assess trustworthy medical diagnostics 1 under clinically grounded confounders.