arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

采样贝尔曼算子的泊松切线极限与关键策略切换

Poisson Tangent Limits and Critical Policy Switching for Sampled Bellman Operators

Ming-Zhe Dai, Chengxi Zhang

arXiv 2608.11549首次发表:更新:

AI 中文总结

本文针对带连续动作空间的折扣马尔可夫决策过程的随机候选动作选择模型,建立了一阶理论,推导了值函数收敛的速率、渐近选择规则及关键策略切换机制,数值实验验证了相关结论。

AI 中文摘要

考虑具有连续动作空间的折扣马尔可夫决策过程,在每次访问状态时,控制器抽取N个候选动作的随机池并从中选择。当最优动作集在采样分布下质量为零时,该随机候选模型的值函数随N增大收敛到最优值,但收敛速率和渐近选择规则由最优集的几何结构与近最优候选动作的转移核共同决定。本文针对这两者建立了精确的一阶理论:重标度后的近最优候选动作收敛到标记泊松点过程,值函数 gap 的主导项是非线性切线贝尔曼算子的不动点,该算子是经典预解式的随机推广,出现于多个最优动作竞争的场景中;精确平局处的渐近选择具有真正的动态性,由转移核通过不动点驱动,且在极限泊松层面上存在显式的 Mecke 积分公式;以临界速率扰动平局会产生一个切换层,其中值函数 gap 和选择在分支 regime 间连续插值;在唯一最优的异质情形下,关键类通过极限策略的折扣可达性以最慢的全局尺度传播。数值实验验证了预测的速率、常数和选择概率。

英文摘要

Consider a discounted Markov decision process with continuous action space in which, at each state visit, the controller draws a random pool of $N$ candidate actions and selects among them. When the optimal action set has zero mass under the sampling distribution, the value of this random-candidate model converges to the optimal value as $N$ grows, but the rate of convergence and the asymptotic selection rule are governed jointly by the geometry of the optimal set and by the transition kernels of the near-optimal candidates. This paper develops an exact first-order theory of both. The rescaled near-optimal candidates converge to a marked Poisson point process, and the leading term of the value gap is the fixed point of a nonlinear tangential Bellman operator, a stochastic generalization of the classical resolvent that emerges when several optimal actions compete. The asymptotic selection at exact ties is genuinely dynamic, driven by the transition kernels through the fixed point, and admits an explicit Mecke integral formula at the limiting Poisson level; perturbing the tie at the critical rate yields a switching layer in which the value gap and the selection interpolate continuously between the branch regimes, and in the unique-optimum heterogeneous case the critical class is propagated at the slowest global scale through the discounted reachability of the limit policy. Numerical experiments confirm the predicted rates, constants, and selection probabilities.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑