通过反事实临床审计揭示ICU脓毒症管理中医务离线强化学习中的毒性模仿行为
Unmasking Toxic Mimicry in Medical Offline Reinforcement Learning for ICU Sepsis Management via Counterfactual Clinical Audits
浏览论文内容
中文总结 AI 辅助
该研究针对ICU脓毒症管理中医务离线强化学习的毒性模仿问题,提出反事实临床审计(CCA)框架,审计发现HCT-RL比MedDT更符合临床指南,证实反事实审计是医疗RL的必要评估标准。
中文摘要 AI 辅助
离线强化学习(RL)为优化ICU治疗决策提供了巨大潜力,但标准评估指标均方误差(MSE)和拟合Q评估(FQE)仅评估行为模仿,无法检测毒性模仿——一种智能体复制有害模式的失效模式,例如在舒适护理过渡期间的治疗弃权(不执行)。我们使用MIMIC-III数据库,提出反事实临床审计(CCA)框架,该框架通过锚定在脓毒症存活运动(SSC)指南中的生理扰动对RL智能体进行压力测试。我们审计了医务决策转换器(MedDT)和历史因果转换器(HCT-RL),其中HCT-RL采用因果动作屏蔽、基于倾向的重要性加权和保守Q学习。CCA显示MedDT在乳酸升高时矛盾地降低血管加压剂剂量,与复苏指南相悖,而HCT-RL维持生理一致的响应。这些发现揭示了统计拟合与临床安全性之间的系统性错位,支持反事实审计作为医疗RL的必要评估标准。
英文摘要
Offline reinforcement learning (RL) offers considerable promise for optimizing ICU treatment decisions, yet standard evaluation metrics Mean Squared Error (MSE) and Fitted Q-Evaluation (FQE) assess only behavioral imitation and cannot detect Toxic Mimicry, a failure mode in which agents replicate harmful patterns such as treatment withdrawal during comfort-care transitions. Using the MIMIC-III database, we propose the Counterfactual Clinical Audit (CCA) framework, which stress-tests RL agents through physiological perturbations anchored in Surviving Sepsis Campaign (SSC) guidelines. We audit a Medical Decision Transformer (MedDT) and a Historical Causal Transformer (HCT-RL), the latter employing Causal Action Shielding, propensity-based importance weighting, and Conservative Q-Learning. CCA reveals that MedDT paradoxically reduces vasopressor dosage as lactate escalates, contradicting resuscitation guidelines, while HCT-RL maintains physiologically consistent responses. These findings expose a systemic misalignment between statistical fit and clinical safety, supporting counterfactual audits as a necessary evaluation standard for medical RL.
发表机构
- School of Engineering, Vanderbilt University(范德堡大学工程学院)
- Pratt School of Engineering, Duke University(杜克大学普拉特工程学院)
机构由 AI 辅助整理,请以论文原文为准。