arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

摊销得分-哈密顿策略迭代:松弛随机控制问题的无网格方案

Amortized Score-Hamiltonian Policy Iteration: A Grid-Free Scheme for Relaxed Stochastic Control Problems

Qi Feng, Gu Wang

arXiv 2610.04285首次发表:更新:

发表机构

Florida State University; Worcester Polytechnic Institute(佛罗里达州立大学; 伍斯特理工学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出摊销无网格的得分-哈密顿策略迭代算法,用于熵正则化松弛随机控制问题,通过共享条件采样器和参数化评论家实现耦合演员-评论家学习,并给出误差界限与高维数值验证。

AI 中文摘要

我们开发了一种摊销的、无网格的连续朗之万动力学实现,用于基于策略-价值迭代的熵正则化、无限时域松弛随机控制问题。精确迭代的改进率是策略与其哈密顿量的吉布斯定律之间的相对费舍尔信息的折扣聚合。相关的得分残差是控制朗之万动力学传输其定律的速度。我们将该速度投影到一个跨状态共享的条件采样器上,而不是每个状态一个朗之万动力学,并将价值动力学投影到一个参数化评论家上,在采样状态处估计两个投影以获得耦合的演员-评论家流。得分损失衡量演员与当前评论家的一致性,而策略评估残差衡量评论家与演员的一致性。我们还推导了梯度和海森残差,包括梯度方程的费曼-卡茨表示,以控制投影价值迭代未检测到的误差。HJB残差的精确分解将这些误差组合成在验证和对数索博列夫假设下的策略次优性界限。在线性二次类中,两个投影都是精确的并恢复逐点迭代,我们提供了通用模型的数值实验,以展示高维中的耦合演员-评论家学习。

英文摘要

We develop an amortized, grid-free implementation of continuous Langevin dynamics based policy-value iteration for entropy-regularized, infinite-horizon relaxed stochastic control problems. The improvement rate of the exact iteration is a discounted aggregate of relative Fisher information between the policy and the Gibbs law of its Hamiltonian. The associated score residual is the velocity with which the control's Langevin dynamics transport its law. We project this velocity onto a conditional sampler shared across states, instead of one Langevin dynamics per state, and the value dynamics onto a parametric critic, estimating both projections at sampled states to obtain coupled actor--critic flows. The score loss measures the actor's agreement with the current critic, while the policy-evaluation residual measures the critic's agreement with the actor. We also derive gradient and Hessian residuals, including a Feynman--Kac representation for the gradient equation, to control errors not detected by the projected value iteration. An exact decomposition of the HJB residual combines these errors into a policy-suboptimality bound under verification and logarithmic Sobolev assumptions. In the linear-quadratic class, both projections are exact and recover the pointwise iteration, and we provide numerical experiments on general models to demonstrate the coupled actor--critic learning in high-dimensions.

Comments25 pages, 5 figures, 3 tables

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑