发表机构
The Chinese University of Hong Kong; University of Toronto(香港中文大学; 多伦多大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究针对通用离散状态空间的连续时间跳跃马尔可夫决策过程,建立q学习理论基础并开发无模型算法,其在网络动态定价应用中优于基准方法。
AI 中文摘要
我们研究具有通用离散状态空间(无需具备向量空间结构)以及连续/离散动作空间的连续时间跳跃马尔可夫决策过程(CTJMDP)中的强化学习(RL)。该设定涵盖了运营领域诸多知名应用,例如带容量约束资源的多产品动态定价(Gallego和van Ryzin,1997)。为建模探索-利用权衡,我们构建了带随机策略的熵正则化连续时间控制问题。近期针对受控扩散过程的连续时间RL技术(Jia和Zhou,2023)聚焦于连续状态空间ℝᵈ,且在理论分析中大量依赖ℝᵈ上的半鞅理论,因此其方法无法直接应用于通用离散状态空间的CTJMDP,这类状态空间可能缺乏欧氏空间固有的代数加减结构。为弥合这一差距,我们建立了CTJMDP下q学习的理论基础,并开发了无模型q学习算法。与朴素时间离散化及用离散时间MDP近似CTJMDP的方法相比,我们的方法具有若干概念和经验优势。在网络动态定价(Gallego和van Ryzin,1997)中的数值实验表明,所提RL算法可可靠学习近最优策略,且始终优于标准基准方法,展现出更优的解质量及对大规模网络实例的有效可扩展性。
英文摘要
We study reinforcement learning (RL) in Continuous-Time Jump Markov Decision Processes (CTJMDPs) featuring general discrete state spaces (which need not possess a vector space structure) and continuous/discrete action spaces. The setup covers many well-known applications in operations such as multi-product dynamic pricing with capacitated resources (Gallego and van Ryzin 1997). To model the exploration-exploitation tradeoff, we formulate an entropy-regularized continuous-time control problem with stochastic policies. Recent continuous-time RL techniques such as $q$-learning for controlled diffusions in (Jia and Zhou 2023) focus on continuous state spaces $\mathbb{R}^d$ and rely heavily on semimartingale theory in $\mathbb{R}^d$ for their theoretical analysis. Consequently, their methods cannot be directly applied to CTJMDPs with general discrete state spaces, which may lack the algebraic addition and subtraction structures inherent to Euclidean spaces. To bridge this gap, we establish the theoretical foundations of $q$-learning for CTJMDPs and develop model-free $q$-learning algorithms. Compared to naïve time discretization and approximating CTJMDPs using discrete-time MDPs, our approach has several conceptual and empirical benefits. Numerical experiments in network dynamic pricing (Gallego and van Ryzin 1997) show that our proposed RL algorithm reliably learns near-optimal policies and consistently outperforms standard benchmark methods, demonstrating superior solution quality and effective scalability to large-scale network instances.