MedRoundsQA:面向多轮医疗咨询的角色与难度感知评估
MedRoundsQA: A Persona and Difficulty Aware Evaluation for Multi-Turn Medical Consultations
- MBZUAI(穆罕默德·本·扎耶德人工智能大学)
- Cairo University(开罗大学)
- Ain Shams University(艾因·夏姆斯大学)
- CSIRO(澳大利亚联邦科学与工业研究组织)
- INSAIT, Sofia University “St. Kliment Ohridski”(索非亚大学INSAIT研究所)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
MedRoundsQA提出一个基于1,387个病例的多轮医疗诊断基准,通过角色化医患对话评估十五个LLM,发现多轮咨询导致诊断性能下降13-39分,且患者教育水平差异造成7-8分的准确率差距,揭示了单轮基准忽视的公平性问题。
AI中文摘要:
医疗基准测试以单轮、多项选择的临床案例为主,无法真实反映实际咨询过程。在实践中,临床医生会交互式地收集证据,且患者的沟通方式差异很大。我们提出了MedRoundsQA,一个基于17个专科的1,387个委员会考试案例构建的多轮诊断基准。每个案例被转换为结构化的24槽位临床记录,然后在不同患者角色下,以底层临床内容保持不变的方式,实例化为受控的医患双智能体对话。我们进一步利用基于模型的不确定性对案例进行难度分类,以支持从易到难的分析。对十五个LLM医生智能体的评估显示:(i)从标准化记录上的单轮诊断转向多轮咨询会导致约13-39分的大幅性能下降;(ii)更多轮次确实能可靠地提高问题相关性,但诊断准确率呈现收益递减,通常在6-12轮后趋于平稳;(iii)患者角色的差异可使诊断准确率产生约7-8分的波动(从最低到最高教育水平),凸显了单轮基准所忽视的公平性风险。
英文摘要:
Medical benchmarks are dominated by single-turn, multiple-choice clinical cases that poorly reflect real consultations. Practically, clinicians elicit evidence interactively and patient communication varies widely. We introduce MedRoundsQA, a multi-turn diagnostic benchmark derived from 1,387 board-exam cases across 17 specialties. Each case is converted into a structured 24-slot clinical record, and then instantiated as controlled doctor-patient dual-agent dialogues under varying patient personas, with the underlying clinical content held fixed. We further classify cases by difficulty using model-based uncertainty to enable easy-to-hard analysis. Evaluations of fifteen LLM doctor agents show that (i) moving from a single-turn diagnosis on the standardized records to multi-turn consultations causes large degradations of roughly 13-39 points; (ii) more turns reliably improves question relevance, but diagnostic accuracy exhibits diminishing returns and typically plateaus after 6-12 turns; and (iii) patient persona differences can shift diagnosis accuracy by about 7-8 points (lowest to highest education), highlighting equity risks that single-turn benchmarks miss.