肿瘤异质性下的自适应化疗控制:基于强化学习的方法
Adaptive Chemotherapy Control under Tumor Heterogeneity via Reinforcement Learning
浏览论文内容
中文总结 AI 辅助
针对肿瘤异质性和耐药性,本研究比较了基于强化学习(TD3与DQN)的闭环化疗给药策略与开环最优控制基准,在100名虚拟患者队列中验证了其有效性,并揭示了疗效与给药一致性之间的权衡。
中文摘要 AI 辅助
设计有效的化疗方案受到肿瘤异质性和耐药性的阻碍,这使针对特定患者的基于模型的最优控制在不同人群中的部署变得复杂。我们开发并比较了在高维异质性肿瘤模型上训练的、具有连续(TD3)和离散(DQN)动作空间的闭环深度强化学习(DRL)给药策略。DRL策略与基于庞特里亚金极大值原理(PMP)推导的开环基准进行了对比。我们使用一个包含100名患者的虚拟队列来评估参数异质性下的泛化能力,该队列在生长和药物敏感性参数上施加正负10%的均匀扰动。在该队列中,TD3实现了更高的平均肿瘤缩减,而DQN则产生了更紧密的患者间给药一致性,揭示了本研究中明显的疗效-一致性权衡。我们的模拟假设对所有肿瘤亚群进行完全观测;向稀疏且有噪声的临床测量的转化将需要部分可观测性公式和/或状态估计。总体而言,结果表明,经过模拟训练的DRL可以学习状态依赖的反馈给药策略,这些策略补充了开环最优控制基准。
英文摘要
Designing effective chemotherapy regimens is hindered by tumor heterogeneity and drug resistance, which complicate the deployment of patient-specific model-based optimal control across diverse populations. We develop and compare closed-loop deep reinforcement learning (DRL) dosing policies with continuous (TD3) and discrete (DQN) action spaces trained on a high-dimensional heterogeneous tumor model. The DRL policies are benchmarked against a Pontryagin's Maximum Principle (PMP)-derived open-loop benchmark. We assess generalization under parametric heterogeneity using a 100-patient virtual cohort with plus or minus 10 percent uniform perturbations in growth and drug-sensitivity parameters. Across this cohort, TD3 achieves higher average tumor reduction, while DQN yields tighter inter-patient dosing consistency, revealing a clear efficacy-consistency trade-off in this study. Our simulations assume full observation of all tumor subpopulations; translation to sparse and noisy clinical measurements will require partial-observability formulations and/or state estimation. Overall, the results show that simulation-trained DRL can learn state-dependent feedback dosing policies that complement open-loop optimal control benchmarks.
发表机构
- University of Texas at Arlington(德克萨斯大学阿灵顿分校)
机构由 AI 辅助整理,请以论文原文为准。