AI 中文总结
本文提出用证据阶梯框架评估医疗保健中的强化学习,强调从回顾性策略到可信干预需逐级验证,并指出历史数据上的高分不等于临床改善,应将其视为社会技术系统中的干预而非单纯优化器。
AI 中文摘要
强化学习(RL)为后果随时间展开的医疗决策提供了一种自然语言,然而大多数已报道的进展仍远未达到常规干预的水平。现有综述按算法或临床应用来组织该领域。我们则通过证据阶梯来审视医疗保健中的强化学习:问题表述、回顾性识别、策略估计、压力测试、前瞻性评估和生命周期监测。这一视角将临床治疗、患者参与和卫生系统运营联系起来,同时揭示了一个反复出现的差距:策略在历史数据集中得分良好,并不能证明它将改善护理。我们综合了每一阶梯上的假设和失败模式,识别了哪些证据可以在不同环境之间转移以及哪些不能,并提出了累积评估的报告实践。静息多臂赌博机被作为一个特例包含在内,而非作为组织框架。核心教训是,医疗保健中的强化学习应被评估为嵌入在不断变化的社会技术系统中的干预措施,而不仅仅是作为回顾性奖励的优化器。
英文摘要
Reinforcement learning (RL) offers a natural language for healthcare decisions whose conse- quences unfold over time, yet most reported progress remains far from routine intervention. Ex- isting surveys organize the field by algorithm or clinical application. We instead review healthcare RL through an evidence ladder: problem formulation, retrospective identification, policy estima- tion, stress testing, prospective evaluation, and lifecycle monitoring. This view connects clinical treatment, patient engagement, and health-system operations while exposing a recurring gap: evi- dence that a policy scores well in a historical dataset is not evidence that it will improve care. We synthesize the assumptions and failure modes at each rung, identify what evidence can and can- not transfer across settings, and propose reporting practices for cumulative evaluation. Restless bandits are included as one special case, not as the organizing framework. The central lesson is that healthcare RL should be evaluated as an intervention embedded in a changing sociotechnical system, rather than only as an optimizer of a retrospective reward.