arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.23127cs.LG

具有泊松决策时段的连续时间情节式MDP中的可证明高效强化学习

Provably Efficient Reinforcement Learning in Continuous-Time Episodic MDPs with Poisson Decision Epochs

发表机构耶鲁大学 · 麻省理工学院 · 加州大学洛杉矶分校
查看机构详情
  • Yale University(耶鲁大学)
  • Massachusetts Institute of Technology(麻省理工学院)
  • University of California, Los Angeles(加州大学洛杉矶分校)

机构由 AI 辅助整理,请以论文原文为准。

Kenny Guo, Valentio Iverson, Sahan Wijetunga, William Chang

首次发表
浏览论文内容

中文总结 AI 辅助

针对连续时间情节式MDP,在泊松决策时段和Lipschitz平滑假设下,扩展UCRL和Q-learning,证明$\widetilde{O}(T^{2/3})$遗憾界及匹配下界,实现最优速率。

中文摘要 AI 辅助

许多现实世界的强化学习(RL)问题在连续时间中演化,其中决策发生在不规则的、事件驱动的间隔,而非固定的离散步骤。我们研究情节式连续时间马尔可夫决策过程(MDP),其中决策时段由齐次泊松过程控制,且奖励和转移动力学随时间平滑变化。我们考虑每情节固定跳跃次数和固定时间预算(具有随机数量的泊松决策时段)两种情况。在时间上的Lipschitz连续性假设下,我们通过离散化利用局部平滑性,并将UCRL(Auer和Ortner 2006)和Q-learning(Jin等人2018)扩展到该设置,为基于模型和无模型算法证明了$\widetilde{O}(T^{2/3})$的遗憾界。最后,我们建立了匹配的$\widetilde{\Omega}(T^{2/3})$极小极大下界,表明该速率在对数因子意义下是最优的。这些结果为具有泊松决策时段的Lipschitz平滑连续时间情节式MDP提供了首个紧致遗憾保证。

英文摘要

Many real-world reinforcement learning (RL) problems evolve in continuous time, where decisions occur at irregular, event-driven intervals rather than at fixed discrete steps. We study episodic continuous-time Markov Decision Processes (MDPs) in which decision epochs are governed by a homogeneous Poisson process and the reward and transition dynamics vary smoothly over time. We consider both a fixed number of jumps per episode and a fixed time budget with a random number of Poisson decision epochs. Under a Lipschitz continuity assumption in time, we exploit local smoothness through discretization and extend both UCRL (Auer and Ortner 2006) and Q-learning (Jin et al. 2018) to this setting, proving $\widetilde{O}(T^{2/3})$ regret bounds for both model-based and model-free algorithms. Finally, we establish matching $\widetildeΩ(T^{2/3})$ minimax lower bounds, showing that the rate is optimal up to logarithmic factors. These results provide the first tight regret guarantees for Lipschitz-smooth continuous-time episodic MDPs with Poisson decision epochs.

补充信息

↑