arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

状态约束策略迭代的惩罚均匀局部化

Penalty-Uniform Localization for State-Constrained Policy Iteration

Yeongjong Kim, Jiwoong Jang, Yeoneung Kim

arXiv 2609.16829首次发表:更新:

发表机构

POSTECH; University of Maryland; Yonsei University(浦项工科大学; 马里兰大学; 延世大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文针对状态约束最优控制中的惩罚奇异局部化问题,证明惩罚自身提供约束抵消增长,在耦合细化条件下给出神经值逼近收敛条件,并通过高维基准和导航实验验证方法。

AI 中文摘要

对状态约束进行惩罚会产生奇异局部化问题:全空间值可能随惩罚参数的倒数增长,而数值扩散甚至允许内向反馈穿越边界。对于确定性折扣最优控制,我们证明了惩罚本身提供的约束可以抵消这种增长。对于具有消失粘性的单调中心差分格式,在网格尺寸$h$和惩罚参数$\varepsilon$下,内向势垒将数值泄漏的额外成本限制为$O(h/\varepsilon)$。在$h\le\varepsilon$条件下,占据估计随后产生$O(h)$的局部化误差,且所需的盒子边距与$1/h$成对数关系,并与$\varepsilon$无关。我们将结果扩展到有界平方距离惩罚,并将其与基于残差的评估界和折扣策略误差传播相结合。在惩罚和离散化误差同时消失的耦合细化下,关于评估误差和贪心间隙的显式条件确保神经值逼近收敛到约束值,允许可测的非唯一贪心选择器。参考计算在网格和惩罚尺度上隔离泄漏和局部化,并区分评估、迭代和逼近误差。显式圆柱状态约束解提供了任意维度的基准,测试至维度二十。成对的障碍物导航实验说明了在比较原始残差和有限网格辅助神经评估时,必须将策略值精度与采样残差一起评估。

英文摘要

Penalizing a state constraint creates a singular localization problem: whole-space values may grow like the inverse penalty parameter, while numerical diffusion allows even an inward feedback to cross the boundary. For deterministic discounted optimal control, we show how the penalty itself supplies confinement that offsets this growth. With mesh size $h$ and penalty parameter $\varepsilon$, an inward barrier bounds the additional cost of numerical leakage by $O(h/\varepsilon)$ for a monotone centered-difference scheme with vanishing viscosity. Under $h\le\varepsilon$, an occupation estimate then yields $O(h)$ localization error with a sufficient box margin logarithmic in $1/h$ and independent of $\varepsilon$. We extend the result to bounded squared-distance penalties and combine it with residual-based evaluation bounds and discounted policy-error propagation. Under coupled refinement with vanishing penalization and discretization errors, explicit conditions on evaluation errors and greedy gaps ensure convergence of the neural value approximations to the constrained value, allowing measurable, nonunique greedy selectors. Reference calculations isolate leakage and localization across mesh and penalty scales, and distinguish evaluation, iteration, and approximation errors. An explicit cylindrical state-constraint solution provides a benchmark in arbitrary dimension, tested up to dimension twenty. Paired obstacle-navigation experiments illustrate why policy-value accuracy must be assessed alongside sampled residuals when comparing raw-residual and finite-grid-assisted neural evaluation.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑