面向慢性多重病症的患者中心治疗规划:用于偏好建模的分层强化学习框架
Patient-Centered Treatment Planning for Chronic Multimorbidity: A Hierarchical Reinforcement Learning Framework for Preference Modeling
浏览论文内容
中文总结 AI 辅助
提出FAHOC分层强化学习框架,通过因子化动作与偏好约束建模患者偏好,在5万共病患者数据上实现QALY改善0.669且不违反偏好。
中文摘要 AI 辅助
患者偏好,定义为患者对临床建议表现出的依从意愿和能力,是治疗效果的主要决定因素,然而在现有的计算治疗规划模型中却结构性缺失。我们通过提出患者中心的因子化动作分层选项-评论家(FAHOC)框架来解决这一空白,该框架是一个分层强化学习(HRL)框架,联合学习对应于治疗策略的高层选项以及将联合动作空间分解为疾病特异性和干预特异性子组件的因子化选项内策略,同时施加合作感知的动作掩蔽机制。这实现了结构化探索、跨层级改进的信用分配以及更具可解释性的决策路径,同时强制执行患者的偏好。形式化保证确立了合作患者比非合作患者获得更高的最优期望健康结果,且因子化Q函数逼近误差被证明是有界的。该框架使用来自美国东南部五家医院约50,000名高血压和2型糖尿病共病患者的纵向数据进行评估。FAHOC实现了质量调整生命年期望等效改善0.669(对比观察到的临床实践为-0.133),在95.9%的病例中正确识别合作患者,并且在保留测试中从未违反患者偏好,表明具有显式偏好约束的HRL能够支持多重病症管理中偏好一致且临床安全的决策。
英文摘要
Patient preference, defined as a patient's demonstrated willingness and capacity to adhere to clinical recommendations, is a primary determinant of therapeutic effect yet remains structurally absent from existing computational treatment planning models. We address this gap by presenting patient-centered factored-action hierarchical option-critic (FAHOC), a hierarchical reinforcement learning (HRL) framework that jointly learns high-level options corresponding to therapeutic strategies and factored intra-option policies that decompose the joint action space into disease- and intervention-specific subcomponents, while imposing a cooperation-aware action masking mechanism. This enables structured exploration, improved credit assignment across hierarchy levels, and more interpretable decision pathways, while enforcing patients' preferences. Formal guarantees establish that cooperative patients achieve higher optimal expected health outcomes than non-cooperative patients, and that the factored Q-function approximation error is provably bounded. The framework is evaluated using longitudinal data collected from approximately 50,000 comorbid hypertension and type 2 diabetes mellitus patients from five hospitals in the Southeast U.S. FAHOC achieves a quality-adjusted life year expectancy equivalent improvement of 0.669 (vs -0.133 observed clinician practice), correctly identifies cooperative patients in 95.9% of cases and never violates a patient's preference in held-out test, demonstrating that HRL with explicit preference constraints can support preference-consistent, clinically safe decision-making in multimorbidity management.
发表机构
- University of Tennessee, Knoxville(田纳西大学诺克斯维尔分校)
- University of Tennessee Health Science Center College of Medicine(田纳西大学健康科学中心医学院)
- University of Tennessee Medical Center(田纳西大学医学中心)
- University of Arkansas(阿肯色大学)
机构由 AI 辅助整理,请以论文原文为准。