arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.18999cs.CLcs.AI

MedDDC-Eval:多轮医疗咨询代理的诊断解耦评估

MedDDC-Eval: Diagnosis-Decoupled Evaluation of Multi-Turn Medical Consultation Agents

  • Baidu, Inc(百度公司)

机构由 AI 辅助整理,请以论文原文为准。

Guofeng Zhang, Yizeng Quan, Huaiyi Fang, Jianwei Lv, Jinyao Liu, Xunxu Duan, Lening An, Yu Ouyang, Junfeng Wang

AI总结:

研究多轮医疗咨询代理评估问题,提出诊断解耦测试平台MedDDC-Eval,通过共享冻结阅读器保持病史到诊断映射不变,能测量诊断有用性等,支持可控归因与策略开发,应用GRPO后训练模型可提升性能。

AI中文摘要:

多轮医疗咨询代理必须决定询问内容、适应患者回答并确定收集的证据何时足够。然而,耦合评估将策略引出的病史质量与特定于策略的终端诊断生成混为一谈:强大的诊断生成可以弥补薄弱的病史,而较弱的诊断生成则可能掩盖丰富的病史。我们引入了MedDDC-Eval,这是一个诊断解耦的测试平台,它将引出的病史视为比较对象,并通过共享的冻结阅读器保持病史到诊断的映射不变。在两个保留源(一个有根据的接口和一个可审计的诊断轨迹效率(D/T/E)工具)上,测量诊断有用性、信息获取和效率。定向语义覆盖随后进行确定性一对一分配,为开放式项目产生连贯的精确召回计数,每个预测或参考最多有一个认可匹配。在保持病史不变的情况下,仅改变诊断阅读器会使诊断F1变化2.2 - 19.0分,并在Record和Dialogue分割上反转18%和36%的成对策略排序。我们进一步在交互式多轮展开中应用标准的组相对策略优化(GRPO),使用诊断结果和轨迹反馈对Qwen3-32B进行后训练。在100例Record和70例Dialogue分割上,训练后的策略比其初始化提高了9.7和4.6个总分点;去除任何一个主要信号都会降低保留的联合性能。这些结果表明,MedDDC-Eval支持可控归因、可解释的引出病史测量以及评估引导的证据获取策略开发。

英文摘要:

Evaluating multi-turn medical consultation agents requires judging the diagnostic support provided by the histories they elicit through interaction. Yet coupled evaluation lets each policy both elicit the history and generate the terminal diagnosis, so a diagnosis score confounds the elicited history with the policy's own terminal diagnosis generator. We introduce MedDDC-Eval, a diagnosis-decoupled evaluation testbed over held-out cases derived from medical records and online consultations. It applies the same frozen shared diagnostic reader to every policy-elicited history, holding terminal diagnosis generation fixed across policies and enabling comparison under the shared diagnostic reader. It reports diagnostic support, information-acquisition coverage, and efficiency. LLM-assisted semantic matching followed by deterministic one-to-one assignment makes the diagnosis-trajectory-efficiency (D/T/E) scores auditable. In a fixed-history audit across eight policies, replacing each policy's own generator with the shared diagnostic reader shifts diagnosis F1 by 2.2-19.0 points and reverses 18% and 36% of pairwise orderings on the Record and Dialogue splits. To examine downstream utility, we use standard Group Relative Policy Optimization (GRPO) with a separate training-time reward that targets the same diagnosis and trajectory dimensions. Relative to its Qwen3-32B initialization, the trained policy gains 9.6 and 4.6 aggregate-score points on the held-out Record and Dialogue splits, respectively, and ablating either feedback signal reduces the aggregate score on both. Together, MedDDC-Eval supports comparison under a shared diagnostic reader and evaluation-informed policy development, while complementing end-to-end evaluation when terminal diagnosis generation is also part of the target capability.

补充信息

↑