AI 中文总结
研究针对膀胱癌复发性疾病治疗需个性化决策的问题,提出集成预测状态转换建模、马尔可夫决策过程和深度Q网络强化学习环境的框架,能动态制定治疗计划,经评估效果良好,凸显其作为个性化膀胱癌治疗规划决策支持框架的潜力。
AI 中文摘要
膀胱癌治疗需要个性化和适应性决策,尤其是复发性疾病,其治疗效果在连续临床阶段会变化。传统临床决策支持系统依赖静态指南或单步预测模型,难以捕捉疾病进展。本文提出用于膀胱癌治疗规划的循环患者状态转换模拟框架,集成预测状态转换建模、马尔可夫决策过程和深度Q网络强化学习环境。预测模块估计治疗后肿瘤特征变化,强化学习智能体通过与模拟患者轨迹交互优化治疗决策。该框架能动态、针对患者制定治疗计划,生成可解释的治疗轨迹和模拟日志,提高透明度并支持临床决策。通过与现有基于强化学习的治疗规划方法对比评估,该框架取得了累积奖励63,918.87、每轮平均训练损失0.0056和策略改进得分6.62%,证明了在模拟复发性治疗环境中的有效序列学习和强大的治疗优化能力。这些发现凸显了强化学习的循环患者状态转换模拟作为个性化膀胱癌治疗规划和人工智能辅助精准肿瘤学灵活决策支持框架的潜力。
英文摘要
Bladder cancer treatment requires personalized and adaptive decision-making, particularly for recurrent disease, where treatment effectiveness changes across successive clinical episodes. Conventional clinical decision support systems typically rely on static treatment guidelines or single-step predictive models, limiting their ability to capture disease progression over time. This paper presents a recurrent patient state-transition simulation framework for bladder cancer treatment planning that integrates predictive state-transition modeling with a Markov Decision Process (MDP) and a Deep Q-Network (DQN) reinforcement learning environment. The predictive module estimates changes in tumor characteristics following treatment, while the reinforcement learning agent sequentially optimizes treatment decisions by interacting with simulated patient trajectories. This framework enables dynamic, patient-specific treatment planning by continuously adapting recommendations to evolving clinical states. It also generates interpretable treatment trajectories and detailed simulation logs to improve transparency and support clinical decision-making. The proposed framework was evaluated against existing reinforcement learning-based treatment planning approaches. It achieved a cumulative reward of 63,918.87, an average training loss per episode of 0.0056, and a policy improvement score of 6.62%, demonstrating effective sequential learning and robust treatment optimization in a simulated recurrent treatment environment. These findings highlight the potential of recurrent patient state-transition simulation with reinforcement learning as a flexible decision-support framework for personalized bladder cancer treatment planning and AI-assisted precision oncology.