发表机构
Instituto Universitario de Tecnologías de la Información y Comunicaciones; Universitat Politècnica de València(信息技术与通信大学学院; 瓦伦西亚理工大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究基于MIMIC-IV数据构建脓毒症血流动力学管理的离线强化学习策略,通过双离线策略评估验证其优于临床实践,为脓毒症决策支持提供新方向。
AI 中文摘要
脓毒症患者的静脉输液和血管加压药给药是在不确定性下做出的序贯决策,主要由临床判断指导,因此成为从历史诊疗数据中应用强化学习的天然目标。由于学习到的策略无法在患者身上进行试验,其价值必须通过离线策略估计得出,而这类估计可能存在脆弱性和乐观性。本研究通过在透明验证框架中结合离线策略估计、可靠性诊断和临床医生一致性分析,推进了脓毒症治疗策略的可靠评估。我们从MIMIC-IV重症监护数据库中选取36872例脓毒症ICU住院患者队列,将输液和血管加压药给药建模为包含1000个状态和25个动作的离散马尔可夫决策过程,状态和动作由输液与血管加压药水平的5×5网格定义,并通过策略迭代求解。临床医生的行为策略采用随机森林估计,缓解了有效样本量(ESS)的崩溃(ESS为50.1,而采用平滑计数时为4.0),否则会使重要性采样估计不稳定。采用加权重要性采样(WIS)和拟合Q评估(FQE)两种估计器对学习到的策略进行评估,以ESS和临床医生一致性作为可靠性检查。经验变量选择发现,状态的组成比其大小更重要。两种估计器均显示学习到的策略回报高于临床医生的回报(WIS为50.8,FQE为46.8,而临床医生为38.2,ESS为50.1),但该策略与观察到的诊疗实践差异较小(总变差为0.18),倾向于减少静脉输液量。这些回顾性单中心离线策略结果表明,学习到的策略是对现有诊疗实践的临床可行改进,为其作为基于不一致性的临床决策支持方法的进一步评估提供了依据。
英文摘要
The dosing of intravenous fluids and vasopressors in sepsis is a sequential decision made under uncertainty and guided largely by clinical judgment, which makes it a natural target for reinforcement learning from historical care. Because a learned policy cannot be trialed on patients, its value must be estimated off-policy, and such estimates can be fragile and optimistic. This work advances the reliable evaluation of sepsis treatment policies by combining off-policy estimation, reliability diagnostics, and clinician-agreement analyses in a transparent validation framework. We modeled fluid and vasopressor dosing on a cohort of 36,872 septic ICU stays drawn from the MIMIC-IV critical-care database, as a discretized Markov decision process with 1,000 states and 25 actions, defined by a five-by-five grid of fluid and vasopressor levels and solved by policy iteration. The clinicians' behavior policy was estimated with a random forest, which mitigated the collapse of the Effective Sample Size (ESS 50.1 against 4.0 with smoothed counts) that otherwise destabilizes the importance-sampling estimate. The learned policy was evaluated with two estimators, weighted importance sampling (WIS) and fitted Q evaluation (FQE), with the ESS and clinician agreement as reliability checks. An empirical variable selection found that the composition of the state matters more than its size. Both estimators place the learned policy above the clinicians' return (WIS 50.8 and FQE 46.8 against 38.2, ESS 50.1), yet it departs only modestly from observed practice (total variation 0.18), favoring less intravenous fluid. These retrospective single-center off-policy results support the learned policy as a clinically plausible refinement of observed practice and motivate its further evaluation as a discordance-based clinical decision-support approach.