arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.12851cs.AI

MedRoundsQA:面向多轮医疗咨询的角色与难度感知评估

MedRoundsQA: A Persona and Difficulty Aware Evaluation for Multi-Turn Medical Consultations

  • MBZUAI(穆罕默德·本·扎耶德人工智能大学)
  • Cairo University(开罗大学)
  • Ain Shams University(艾因·夏姆斯大学)
  • CSIRO(澳大利亚联邦科学与工业研究组织)
  • INSAIT, Sofia University “St. Kliment Ohridski”(索非亚大学INSAIT研究所)

机构由 AI 辅助整理,请以论文原文为准。

Youssef Mohamed, Ahmed Heakl, Qinrong Cui, Junhong Liang, Rafiq Ali, Bdour Babillie, Nazira Dunbayeva, Lang Gao, Omar Hussein, Ahmed Nada, Ahmed Mohamed Magdy M… 展开作者

Youssef Mohamed, Ahmed Heakl, Qinrong Cui, Junhong Liang, Rafiq Ali, Bdour Babillie, Nazira Dunbayeva, Lang Gao, Omar Hussein, Ahmed Nada, Ahmed Mohamed Magdy Mohamed, Jinghui Liu, Salman Khan, Imran Razzak, Yuxia Wang, Xiuying Chen

AI总结:

MedRoundsQA提出一个基于1,387个病例的多轮医疗诊断基准,通过角色化医患对话评估十五个LLM,发现多轮咨询导致诊断性能下降13-39分,且患者教育水平差异造成7-8分的准确率差距,揭示了单轮基准忽视的公平性问题。

AI中文摘要:

医疗基准测试以单轮、多项选择的临床案例为主,无法真实反映实际咨询过程。在实践中,临床医生会交互式地收集证据,且患者的沟通方式差异很大。我们提出了MedRoundsQA,一个基于17个专科的1,387个委员会考试案例构建的多轮诊断基准。每个案例被转换为结构化的24槽位临床记录,然后在不同患者角色下,以底层临床内容保持不变的方式,实例化为受控的医患双智能体对话。我们进一步利用基于模型的不确定性对案例进行难度分类,以支持从易到难的分析。对十五个LLM医生智能体的评估显示:(i)从标准化记录上的单轮诊断转向多轮咨询会导致约13-39分的大幅性能下降;(ii)更多轮次确实能可靠地提高问题相关性,但诊断准确率呈现收益递减,通常在6-12轮后趋于平稳;(iii)患者角色的差异可使诊断准确率产生约7-8分的波动(从最低到最高教育水平),凸显了单轮基准所忽视的公平性风险。

英文摘要:

Medical benchmarks are dominated by single-turn, multiple-choice clinical cases that poorly reflect real consultations. Practically, clinicians elicit evidence interactively and patient communication varies widely. We introduce MedRoundsQA, a multi-turn diagnostic benchmark derived from 1,387 board-exam cases across 17 specialties. Each case is converted into a structured 24-slot clinical record, and then instantiated as controlled doctor-patient dual-agent dialogues under varying patient personas, with the underlying clinical content held fixed. We further classify cases by difficulty using model-based uncertainty to enable easy-to-hard analysis. Evaluations of fifteen LLM doctor agents show that (i) moving from a single-turn diagnosis on the standardized records to multi-turn consultations causes large degradations of roughly 13-39 points; (ii) more turns reliably improves question relevance, but diagnostic accuracy exhibits diminishing returns and typically plateaus after 6-12 turns; and (iii) patient persona differences can shift diagnosis accuracy by about 7-8 points (lowest to highest education), highlighting equity risks that single-turn benchmarks miss.

↑