arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

超立方体状态空间的边界感知强化学习:基于确定性策略梯度

Boundary-aware Reinforcement Learning for Hypercube State Spaces via Deterministic Policy Gradient

Lijun Bo, Yijie Huang, Chenhao Lu

arXiv 2610.09712首次发表:更新:

发表机构

Xidian University; The Hong Kong Polytechnic University; University of Science and Technology of China(西安电子科技大学; 香港理工大学; 中国科学技术大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对超立方体反射状态动力学的强化学习,提出连续时间确定性策略梯度框架,通过诺伊曼边界条件感知方法降低边界残差并提升学习稳定性,在水库控制实验中验证有效性。

AI 中文摘要

我们为具有反射状态动力学的强化学习开发了一个连续时间确定性策略梯度框架,其中状态过程由超立方体上的受控反射随机微分方程支配。在适当的正则性假设下,我们建立了价值函数与诺伊曼贝尔曼方程之间的联系,引入了一个优势率函数,该函数产生确定性策略梯度公式,并证明了鞅刻画定理。受这些理论结果的启发,我们提出了一种用于反射随机系统的连续时间深度确定性策略梯度算法,其中诺伊曼边界条件通过软惩罚或硬架构约束来施加。我们进一步量化了理想连续时间动力学与实际执行中离散采样的探索动力学之间的差异,表明随着时间网格细化且探索噪声消失,误差会衰减。我们在水库控制问题上的实验展示了该强化学习框架的有效性,强调了边界感知方法显著减少了诺伊曼边界残差并增强了学习稳定性。

英文摘要

We develop a continuous-time deterministic policy gradient framework for reinforcement learning with reflected state dynamics, where the state process is governed by a controlled reflected stochastic differential equation on a hypercube. Under suitable regularity assumptions, we establish the connection between the value function and the Neumann Bellman equation, introduce an advantage-rate function that yields a deterministic policy gradient formula, and prove the martingale characterization theorem. Motivated by these theoretical results, we propose a continuous-time deep deterministic policy gradient algorithm for reflected stochastic systems, in which the Neumann boundary condition is imposed via either soft penalization or hard architectural constraint. We further quantify the discrepancy between the ideal continuous-time dynamics and the discretely sampled exploratory dynamics executed in practice, showing that the error decays as the time grid is refined and exploration noise vanishes. Our experiments on reservoir control problems illustrate the effectiveness of the RL framework, highlighting that boundary-aware methods substantially reduce Neumann boundary residuals and enhance learning stability.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑