发表机构
Peking University; Northeastern University; School of Basic Medical Sciences, Peking University(北京大学; 东北大学; 北京大学基础医学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究针对主动医疗对话提出ESR严重度感知问题监督目标,在DDxPlus数据集上提升了高严重度诊断准确率、降低了漏诊率,推动主动医疗对话向感知后果的证据获取方向发展。
AI 中文摘要
主动医疗对话要求智能体基于不完整的患者信息决定需要询问的内容。现有的信息获取方法通常优先选择最能降低诊断不确定性的问题,虽然这些方法在获取有价值的证据方面有效,但该标准忽略了医疗诊断的一个重要特性:不同诊断错误可能带来截然不同的后果。遗漏严重疾病的影响远大于降低低风险替代方案的不确定性。因此,问题获取不仅应考虑新证据的信息量,还应考虑其对下游诊断决策的预期影响。为此,我们提出了预期严重度风险(Expected-Severity-Risk, ESR),这是一种感知后果的问题监督目标,通过候选问题预期降低的严重度感知终端风险来评估其价值。由于必须在观察到答案前选择问题,ESR 使用仅训练时的总体统计数据对可能的答案进行边缘化处理,随后将其排名提炼为仅基于前缀的语言策略,因此在部署时选择下一个问题无需教师端计算。在 DDxPlus 数据集上使用 3 个 Qwen3-4B 训练种子进行实验,匹配 ESR 监督将平均高严重度诊断遗漏率从 0.0645 降至 0.0455(下降 29.5%),平均诊断准确率从 0.9123 提升至 0.9320,且每对话仅需增加 0.14 个问题。固定预算分析显示,当控制问题数量时,两个目标的行为仍存在差异;而匹配的预期 0/1 风险对照实验表明,严重度感知加权相比通用决策感知监督,能更好地改善高严重度错误情况。这些结果支持主动医疗对话从单纯降低不确定性转向感知后果的证据获取。
英文摘要
Proactive medical dialogue requires an agent to decide what to ask from incomplete patient information. Existing information-seeking approaches commonly prioritize questions that most reduce diagnostic uncertainty, but this criterion overlooks an important property of medical diagnosis: different diagnostic errors can carry substantially different consequences. The most informative question may therefore differ from the one most valuable for the downstream decision. We propose Expected-Severity-Risk (ESR), a consequence-aware question-supervision objective that values each candidate by its expected reduction in severity-aware terminal risk. Because questions must be selected before their answers are observed, ESR marginalizes over possible answers using train-only population statistics. Its rankings are then distilled into a prefix-only language policy, requiring no teacher-side risk computation at deployment. Across three matched Qwen3-4B training seeds on DDxPlus, ESR reduces mean high-severity diagnostic miss from 0.0645 to 0.0455 (29.5% relative reduction) and improves mean diagnostic accuracy from 0.9123 to 0.9320 while requiring only 0.14 additional questions per dialogue. Fixed-budget analyses show that the distinction persists when question count is controlled, while a matched expected-0/1-risk student control further isolates the contribution of asymmetric severity weighting. These results support moving proactive medical dialogue beyond uncertainty reduction toward consequence-aware evidence acquisition.
Comments5 pages, 2 figures, 2 tables. Submitted to ICASSP 2027. Updated author list and manuscript