MedGym:面向动态医疗治疗强化学习的统一连续时间基准
A Unified Benchmark for Dynamic Medical Treatment Reinforcement Learning
- Tokyo University of Agriculture and Technology(东京农业大学)
- Institute of Science Tokyo(东京科学研究院)
- National University of Singapore(国立新加坡大学)
- LY Corporation(LY公司)
- Altos Labs, Inc.(Altos实验室)
- National Institute of Advanced Industrial Science and Technology (AIST)(国家先进工业科学与技术研究院)
- Emory University(埃默里大学)
- Norwegian University of Science and Technology(挪威科学技术大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
提出MedGym基准,通过连续时间框架和物理信息神经网络构建可配置的医疗RL环境,支持离散与连续时间方法在非规则治疗间隔下的比较,并评估个性化、轨迹安全等临床指标。
AI中文摘要:
医疗治疗推荐给强化学习(RL)带来了若干挑战:患者生理状态在连续时间内演变,测量和干预以不规则间隔进行,且治疗效果在不同个体间差异显著。然而,现有的RL公式和模拟环境基于离散时间的MDP或POMDP抽象,具有固定或预先指定的决策间隔。因此,评估RL方法能否处理时间间隔依赖的疾病进展、个性化治疗反应以及连续测量点之间的安全性仍然困难。为弥补这一空白,我们引入了MedGym,一个用于动态治疗推荐的基准环境。MedGym在连续时间框架中对纵向患者演变进行建模,并通过使用物理信息神经网络从临床数据构建可配置的医疗RL基准。所得基准支持离线RL和在线RL,并能够在非规则治疗时机和患者特定动态下直接比较离散时间与连续时间方法。此外,MedGym支持从临床重要角度进行评估,包括个性化、轨迹级安全性以及基于模型的离线学习与在线部署之间的性能差距。通过为连续时间动态治疗提供标准化且可配置的基准,MedGym旨在促进对医疗RL方法进行更真实、更具信息量的评估。
英文摘要:
Medical treatment recommendation poses several challenges to reinforcement learning (RL): patient physiology evolves in continuous time, measurements and interventions are performed at irregular intervals, and treatment effects vary substantially across individuals. Existing RL formulations and simulated environments, however, are based on discrete-time MDPs with fixed decision intervals. Thus, it remains difficult to evaluate whether RL methods can handle time-interval-dependent disease progression, personalized treatment response, and safety between consecutive measurement points. To address this gap, we introduce MedGym, a benchmark environment for dynamic treatment recommendation. MedGym models longitudinal patient evolution in a continuous-time framework and constructs a configurable medical RL benchmark from clinical data by using Physics-Informed Neural Networks. The resulting benchmark enables direct comparison between discrete-time and continuous-time methods under irregular treatment timing and patient-specific dynamics. Furthermore, MedGym supports evaluation from clinically important perspectives, such as personalization and trajectory-level safety. By providing a standardized and configurable benchmark for continuous-time dynamic treatment, MedGym enables more realistic and informative evaluation of medical RL methods.