arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2601.06636cs.CLcs.AI

MedEinst:通过反事实差异诊断基准测试医疗大语言模型中的Einstellung效应

MedEinst: Benchmarking the Einstellung Effect in Medical LLMs through Counterfactual Differential Diagnosis

  • Stanford University(斯坦福大学)
  • Xi’an Jiaotong University(西安交通大学)
  • Shenzhen University(深圳大学)
  • Renmin University of China(中国人民大学)

机构由 AI 辅助整理,请以论文原文为准。

Wenting Chen, Zhongrui Zhu, Guolin Huang, Wenxuan Wang

更新

AI总结:

MedEinst通过反事实差异诊断基准测试,揭示医疗大语言模型中的Einstellung效应,并提出ECR-Agent提升循证医学标准。

AI中文摘要:

尽管在医疗基准测试中实现了高准确性,LLMs在临床诊断中表现出Einstellung效应——依赖统计捷径而不是患者特异性证据,在非典型病例中导致误诊。现有基准测试未能检测到这种关键故障模式。我们引入了MedEinst,一个包含5383对临床病例的反事实基准测试,涵盖49种疾病。每对病例包含一个对照病例和一个“陷阱”病例,其区分性证据被修改,从而改变诊断。我们通过偏见陷阱率(误诊陷阱的概率,即使正确诊断了对照病例)来衡量易受性。对17个LLM的广泛评估显示,前沿模型虽然具有高基线准确性,但存在严重的偏见陷阱率。因此,我们提出了ECR-Agent,通过两个组件对齐LLM推理与循证医学标准:(1)动态因果推理(DCI)通过双路径感知、跨三个层次(关联、干预、反事实)的动态因果图推理以及证据审计进行结构化推理;(2)批评驱动的图和记忆演变(CGME)通过在示例库中存储已验证的推理路径并整合疾病特异性知识到演进的疾病图中,不断优化系统。源代码将被发布。

英文摘要:

Despite achieving high accuracy on medical benchmarks, LLMs exhibit the Einstellung Effect in clinical diagnosis--relying on statistical shortcuts rather than patient-specific evidence, causing misdiagnosis in atypical cases. Existing benchmarks fail to detect this critical failure mode. We introduce MedEinst, a counterfactual benchmark with 5,383 paired clinical cases across 49 diseases. Each pair contains a control case and a "trap" case with altered discriminative evidence that flips the diagnosis. We measure susceptibility via Bias Trap Rate--probability of misdiagnosing traps despite correctly diagnosing controls. Extensive Evaluation of 17 LLMs shows frontier models achieve high baseline accuracy but severe bias trap rates. Thus, we propose ECR-Agent, aligning LLM reasoning with Evidence-Based Medicine standard via two components: (1) Dynamic Causal Inference (DCI) performs structured reasoning through dual-pathway perception, dynamic causal graph reasoning across three levels (association, intervention, counterfactual), and evidence audit for final diagnosis; (2) Critic-Driven Graph and Memory Evolution (CGME) iteratively refines the system by storing validated reasoning paths in an exemplar base and consolidating disease-specific knowledge into evolving illness graphs. Source code is to be released.

补充信息

↑