arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.29582cs.CLcs.AI

SUP-MIMIC:用于评估大型语言模型对矛盾证据鲁棒性的多任务临床诊断基准

SUP-MIMIC: A Multi-Task Clinical Diagnosis Benchmark for Evaluating LLMs' Robustness to Contradictory Evidence

  • Beijing University of Technology(北京工业大学)
  • Beijing Institute of Technology(北京理工大学)
  • China-Japan Friendship Hospital(中日友好医院)
  • Dongbei University of Finance and Economics(东北财经大学)

机构由 AI 辅助整理,请以论文原文为准。

Yi Yu, Bo Wang, Chong Feng, Ge Shi, Xia Liu, Ziyi Yang, Xuewen Shi

AI总结:

本研究提出多任务临床诊断基准SUP-MIMIC,评估发现当前先进LLMs在诊断分歧与聚合任务上性能大幅下降,存在漏诊风险,为提升医疗场景下LLMs安全性提供了方法与路线图。

AI中文摘要:

当前对大型语言模型(LLMs)的评估主要聚焦于事实知识检索,却忽略了应对临床指标与诊断间复杂非双射映射这一根本挑战。现有基准无法评估大型语言模型是否真正具备诊断歧义场景所需的推理能力——即相同临床表现可能对应不同病因的情况,也无法评估其诊断聚合场景的能力——即不同症状最终指向同一疾病的情况。为解决该问题,我们提出SUP-MIMIC,这是一个基于MIMIC-IV-v3.1的多任务框架,包含基础评估(BA)、诊断分歧任务(DDT)和诊断聚合任务(DCT)。其中,DDT旨在评估模型在表型相似病例间的“一对多”歧义消除能力,而DCT则评估模型识别不同病理生理通路间“多对一”诊断模式的能力。对当前最先进LLMs的综合评估显示,与基线任务相比,其在DDT和DCT上的性能大幅下降,暴露出模型系统性依赖统计捷径而非真正因果推理的问题。我们的发现进一步凸显了模型对“健康”预测的保守偏向,意味着在现实医疗场景中存在漏诊的重大风险。本研究建立了量化临床推理鲁棒性的严谨方法,并为提升语言模型在临床医学中的安全性提供了路线图。

英文摘要:

Current evaluations of large language models (LLMs) primarily focus on factual knowledge retrieval, overlooking the fundamental challenge of navigating the complex, non-bijective mappings between clinical indicators and diagnoses. Existing benchmarks fail to assess whether large language models truly possess the reasoning capability required for diagnostic ambiguity scenarios, where identical clinical presentations may correspond to different etiologies, and diagnostic convergence scenarios, where heterogeneous symptoms ultimately indicate the same disease. To address this issue, we propose SUP-MIMIC, a multi-task framework utilizing MIMIC-IV-v3.1 that comprises Basic Assessment (BA), Diagnostic Divergence Task (DDT), and Diagnostic Convergence Task (DCT). Specifically, DDT is designed to evaluate the model's "one-to-many" disambiguation capability among phenotypically similar cases, while DCT assesses the model's ability to identify "many-to-one" diagnostic patterns across different pathophysiological pathways. Comprehensive evaluation of state-of-the-art LLMs reveals substantial performance degradation on DDT and DCT compared to baseline tasks, exposing a systemic reliance on statistical shortcuts over genuine causal reasoning. Our findings further highlight a conservative bias toward "healthy" predictions, implying non-trivial risks for missed diagnoses in realistic medical settings. This work establishes a rigorous methodology for quantifying clinical reasoning robustness and provides a roadmap for enhancing the safety of language models in clinical medicine.

补充信息

↑