多轮医学诊断基准测试:Hold、Lure 和 Self-Correction
Benchmarking Multi-turn Medical Diagnosis: Hold, Lure, and Self-Correction
- The University of Texas at Austin(德克萨斯大学奥斯汀分校)
- New York University(纽约大学)
- The University of Texas Southwestern Medical Center(德克萨斯大学西南医学中心)
- University of North Carolina at Chapel Hill(北卡罗来纳大学教堂山分校)
- University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本文通过MINT基准测试揭示了大语言模型在多轮证据积累中的行为模式,发现提前回答、自我修正和强诱因影响诊断决策,提出延迟提问和保留关键证据可提升诊断准确性。
AI中文摘要:
大型语言模型(LLMs)在单轮提供全部临床信息时能实现高诊断准确性,但其在多轮证据积累下的表现尚不清楚。我们引入MINT(医学增量N轮基准),包含1,035个病例,具有临床标注的证据碎片、可控轮次粒度和信息保留分解。通过系统评估11个LLM,发现三种持续行为模式:(1)回答意图,模型在充分证据前急于回答,超过55%的答案在前两轮提交;(2)自我修正,错误到正确答案的修正率是正确到错误翻转的10.6倍,揭示了自我修正的潜在能力;(3)强诱因,如实验室结果等临床显著信息会触发提前回答,即使模型被明确指示等待。我们将这些发现转化为临床可操作的指导:延迟诊断问题可减少提前回答并提高首次承诺点的准确性,保留关键临床证据可防止因提前承诺导致的23.3%准确性下降。本工作提供了受控评估框架和改进LLM在多轮医学诊断中可靠性的具体建议。
英文摘要:
Large language models (LLMs) achieve high accuracy in medical diagnosis when all clinical information is provided in a single turn, yet how they behave under multi-turn evidence accumulation closer to real clinical reasoning remains unexplored. We introduce MINT (Medical Incremental N-Turn Benchmark), a high-fidelity, multi-turn medical diagnosis benchmark comprising 1,035 cases with clinically labeled evidence shards, controlled turn granularity, and information-preserving decomposition. Through systematic evaluation of 11 LLMs on MINT, we uncover three persistent behavioral patterns that significantly impact diagnostic decisions: (1) intent to answer, models rush to answer before sufficient evidence has been observed, with over 55% of answers committed within the first two turns; (2) self-correction, incorrect-to-correct answer revisions occur at up to 10.6 times the rate of correct-to-incorrect flips, revealing a latent capacity for self-correction that premature commitment forecloses; and (3) strong lures, clinically salient information such as laboratory results trigger premature answering even when models are explicitly instructed to wait. We translate these findings into clinically actionable guidance: deferring the diagnostic question to later turns reduces premature answering and improves accuracy at the first point of commitment by up to 62.6%, while reserving salient clinical evidence for later turns prevents a catastrophic accuracy drop of up to 23.3% caused by premature commitment. Our work provides both a controlled evaluation framework and concrete recommendations for improving the reliability of LLMs in multi-turn medical diagnosis.