PCQC:用于多轮医疗对话的特权反事实问题信用
PCQC: Privileged Counterfactual Question Credit for Multi-Turn Medical Dialogue
- Peking University(北京大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
PCQC利用特权患者信息构建反事实问题答案,通过诊断评分器分配问题信用,在四个医疗基准上以更少询问轮次显著提升诊断准确率。
AI中文摘要:
大型语言模型(LLMs)在医学问答方面取得了实质性进展,然而有效的医疗对话还需要学习提出能够揭示相关患者信息的问题。为了训练此类对话策略,一种常见流程将监督微调与基于最终诊断正确性的强化学习(RL)相结合。然而,这种基于结果的监督并不能直接区分单个问题的贡献,并且对于未执行的替代方案不提供问题级别的反馈。为了解决这一差距,我们引入了PCQC(特权反事实问题信用),它在训练期间利用特权患者信息来学习从未被问及的问题。在训练期间,PCQC通过使用特权患者事实来构建替代问题的答案,使得在同一对话状态下替代问题可以直接比较。一个冻结的诊断评分器通过每个问题-答案对支持正确诊断的强度来评估其诊断效用。PCQC将这些比较转化为相对问题信用,教导策略应优先考虑哪些问题,直接监督已执行和未执行的问题,并与基于结果的RL相结合,而无需对未执行的替代方案进行完整轨迹展开。在四个医学基准上的广泛实验表明,PCQC实现了63.10%的平均诊断准确率,分别比GRPO和ATPO高出4.38和4.21个百分点。这些增益是在比GRPO少33.1%的询问轮次下实现的。
英文摘要:
Large language models (LLMs) have made substantial progress on medical question-answering, yet effective medical dialogue also requires learning to ask questions that uncover relevant patient information. To train such dialogue policies, a common pipeline combines supervised fine-tuning with reinforcement learning (RL) based on final diagnostic correctness. However, this outcome-based supervision does not directly distinguish the contributions of individual questions and provides no question-level feedback for unexecuted alternatives. To address this gap, we introduce PCQC (Privileged Counterfactual Question Credit), which uses privileged patient information during training to learn from questions never asked. During training, PCQC makes alternative questions directly comparable at the same dialogue state by using privileged patient facts to construct their answers. A frozen diagnostic scorer evaluates the diagnostic utility of each resulting question-answer pair by how strongly it supports the correct diagnosis. PCQC turns these comparisons into relative question credit that teaches the policy which questions to favor, directly supervising both executed and unexecuted questions alongside outcome-based RL without requiring complete rollouts for the unexecuted alternatives. Extensive experiments across four medical benchmarks demonstrate that PCQC achieves 63.10% mean diagnostic accuracy, outperforming GRPO and ATPO by 4.38 and 4.21 percentage points, respectively. These gains are achieved with 33.1% fewer inquiry turns than GRPO.