增强患者偏好整合强化学习(APP-RL)用于估计最优动态治疗策略
Augmented Patient Preference Incorporated Reinforcement Learning (APP-RL) to Estimate the Optimal Dynamic Treatment Regime
浏览论文内容
中文总结 AI 辅助
本研究提出APP-RL方法,通过数据增强整合患者潜在偏好到基于树的强化学习中,用于估计多阶段、多治疗的最优动态治疗策略,经模拟验证具有鲁棒性和高效性。
中文摘要 AI 辅助
动态治疗策略(DTRs)是顺序决策规则,根据每位患者在每个治疗阶段的既往临床病程,为其个体化定制治疗方案。现有文献通常考虑每位患者的病史,但忽视了患者的偏好。我们提出一种方法,通过数据增强将患者的潜在偏好整合到基于树的强化学习方法中,以估计多阶段、多治疗设置下的最优动态治疗策略。对于每个阶段的每位患者,我们根据问卷回答推导偏好的后验分布,随后用估计的偏好对多个结果进行加权,以确定最优的阶段个性化决策。对于多阶段情形,我们在每个阶段生长一棵决策树,并使用逆向归纳法递归实现该算法。我们提出的方法,命名为增强患者偏好整合强化学习(APP-RL),具有鲁棒性、高效性,并能产生可解释的DTR估计。通过模拟研究,对所提方法的有限样本性能进行了全面评估。
英文摘要
Dynamic treatment regimes (DTRs) are sequential decision rules that individualize treatments to each patient at each treatment stage adapting to their past clinical course. Existing literature typically accommodates each individual's medical history, but overlooks a patient's preferences. We propose a method that incorporates a patient's latent preferences through data augmentation into a tree-based reinforcement learning method to estimate optimal dynamic treatment regimes for multi-stage, multi-treatment settings. For each patient at each stage, we derive the posterior distribution of preferences given responses to a questionnaire, and then subsequently weight multiple outcomes with the estimated preferences to identify the optimal stage-wise personalized decision. For multiple stage situations, we grow a decision tree at each stage and implement the algorithm recursively using backward induction. Our proposed method, named Augmented Patient Preference incorporated Reinforcement Learning (APP-RL) is robust, efficient, and leads to interpretable DTR estimation. The finite-sample performances of the proposed method has been thoroughly evaluated through simulation studies.