AI 中文总结
该研究针对医疗诊断模型忽略检查成本与价值权衡的问题,提出CDPR方法,将诊断建模为成本感知序列决策过程,集成到GRPO后在多数据集上提升诊断准确率并减少检查数量与成本。
AI 中文摘要
临床诊断是一个逐步进行的、需考虑成本的过程:医生会一次开具一项检查,观察结果后更新诊断,最终得出结论。大多数医疗语言模型却将诊断视为一次性分类任务,忽略检查价值与成本之间的权衡。我们将诊断建模为成本感知的序列决策过程,并使用强化学习训练策略。主要难点在于信用分配:唯一可靠的信号来自长轨迹的终点,因此它会将浪费的检查流程与高效流程同等评分。我们提出CDPR(Counterfactual Diagnostic Process Reward,反事实诊断过程奖励),它无需专家标签,也无需学习型评论家。CDPR首先利用动作分布的不确定性找到策略犹豫不决的状态,然后通过策略自身会考虑的替代方案的优势对所选动作进行评分,该优势通过短rollout(回滚)估算,所用效用函数平衡了正确性、检查数量、成本及不可行请求。rollout缓存会复用批次内轨迹以控制成本。我们将CDPR集成到GRPO中,并在一个域内基准(MIMIC-IV)和两个域外基准(ClinicalBench及一家私立医院数据集)上进行测试。CDPR在提高诊断准确率的同时,显著减少了检查的数量和成本。
英文摘要
Clinical diagnosis is a step-by-step, cost-aware process: a physician orders examinations one at a time, observes the results, and updates the diagnosis before reaching a final conclusion. Most medical language models instead treat diagnosis as a one-pass classification task and ignore the trade-off between a test's value and its cost. We model diagnosis as a cost-aware sequential decision process and train the policy with reinforcement learning. The main difficulty is credit assignment: the only reliable signal comes once at the end of a long trajectory, so it scores a wasteful workup the same as an efficient one. We propose CDPR (Counterfactual Diagnostic Process Reward), which needs no expert labels and no learned critic. CDPR first finds the states where the policy hesitates, using the uncertainty of its action distribution, and then scores the chosen action by its advantage over the alternatives the policy itself would consider, estimated with short rollouts under a utility that balances correctness against test count, cost, and infeasible requests. A rollout cache reuses within-batch trajectories to keep the cost low. We integrate CDPR into GRPO and test it on one in-domain (MIMIC-IV) and two out-of-domain (ClinicalBench and a private hospital dataset) benchmarks. CDPR improves diagnostic accuracy while clearly reducing the number and cost of examinations.