伪装成改进的混杂因素:针对129000例患者登记研究中卒中抗血栓治疗的离线强化学习的系统评估
Confounding Masquerading as Improvement: A Systematic Evaluation of Offline Reinforcement Learning for Stroke Antithrombotic Treatment in a 129,000-Patient Registry
浏览论文内容
中文总结 AI 辅助
本研究在12.9万例卒中患者登记数据中系统评估离线强化学习,发现其策略表观改进实为混杂因素导致,去混杂后无临床意义的总体改进,还提供了评估检查表。
中文摘要 AI 辅助
近期离线强化学习(RL)研究报告称,其策略在临床结局上的表现优于医生决策。我们对来自全国登记库(N=129033)的44894例2018年后急性缺血性卒中患者,开展了5类离线RL算法族与14种奖励设计的系统、部分交叉评估。标准拟合Q评估(FQE)得出的表观策略改进估计值为+0.0069;加入早期神经功能恶化惩罚后,该值升至+0.0101。我们识别出奖励嵌入的混杂因素,其中代理终端奖励同时编码基线严重程度、预后及治疗效果。2×2析因分析发现,终端奖励混杂因素占观测信号变化的218.6%,因此去除该因素会超出零假设。经DML启发的GBM奖励残差化后,FQE估计值衰减至+0.0033(p=0.132),完全去混杂后为+0.0025(p=0.291)。基于FQE的诊断、T-学习者分析及直接复发分析均表明,不存在临床有意义的总体改进。1年mRS析因分析重复了该衰减结果。我们提供了基于经验的六步评估检查表;NIHSS分层异质性为前瞻性试验设计提供假设,医院层面的分歧在完全奖励去混杂后不再持续。
英文摘要
Recent offline reinforcement learning (RL) studies report policies that outperform physician decisions on clinical outcomes. We conduct a systematic, partially crossed evaluation of five offline RL algorithm families and 14 reward designs in 44,894 post-2018 acute ischemic stroke patients from a nationwide registry (N = 129,033). Standard Fitted Q-Evaluation (FQE) yields an apparent policy-improvement estimate of +0.0069; adding an Early Neurological Deterioration penalty increases it to +0.0101. We identify reward-embedded confounding, in which a proxy terminal reward encodes baseline severity and prognosis as well as treatment efficacy. A 2 x 2 factorial analysis finds that terminal reward confounding accounts for 218.6% of the observed signal change, so its removal overshoots the null. After DML-inspired GBM reward residualization, the FQE estimate attenuates to +0.0033 (p = 0.132), and full deconfounding yields +0.0025 (p = 0.291). FQE-based diagnostics, T-learner analyses, and direct recurrence analyses converge away from a clinically meaningful aggregate improvement. A 1-year mRS factorial analysis replicates the attenuation. We provide an empirically motivated six-step evaluation checklist. NIHSS-stratified heterogeneity is hypothesis-generating for prospective trial design; hospital-level disagreement does not persist after full reward deconfounding.
发表机构
- AI Research Center, JLK Inc.(JLK公司AI研究中心)
机构由 AI 辅助整理,请以论文原文为准。