AI 中文总结
该研究针对带熵正则化的时间不一致MDP,通过加权泛函分析等方法证明了正则化与原未正则化均衡的存在性,提出的策略迭代算法在强贴现下指数收敛且可构成加权ε-均衡。
AI 中文摘要
我们研究具有可数无限状态空间和无界奖励函数的无限时域时间不一致马尔可夫决策过程(MDP)。该奖励允许显式依赖于初始时间和初始状态,从而适配各类时间不一致的来源。我们寻求松弛反馈均衡,方法基于熵正则化和加权泛函分析方法。通过熵正则化,我们利用不动点算子刻画正则松弛均衡。引入两个具有不同作用的权重函数,一个控制奖励和值函数的增长,另一个定义环境加权空间,我们在乘积拓扑下构造紧不变集,并应用绍德尔-蒂霍诺夫不动点定理证明正则化均衡的存在性。重要的是,该不变集可针对小熵权重λ∈(0,1]统一选取。随后令λ→0+,我们通过紧性、吉布斯策略的集中性及一致可积性论证,证明其子序列极限是原未正则化问题的松弛均衡。我们进一步研究熵正则化均衡问题的策略迭代算法(PIA)。在加权贴现结构和足够强的贴现条件下,我们在合适的加权巴拿赫空间中建立正则化均衡的指数收敛性与唯一性。结合策略迭代误差与定量软最大近似界,我们证明迭代策略构成原未正则化问题的加权ε-均衡,并推导显式遗憾估计。文中还提供了说明强贴现下PIA收敛的数值例子,以及展示弱贴现下其失效的反例。
英文摘要
We study infinite-horizon time-inconsistent Markov decision processes with a countably infinite state space and unbounded reward functions. The reward is allowed to depend explicitly on the initial time and initial state, thereby accommodating general sources of time inconsistency. We seek relaxed feedback equilibria, and our approach is based on entropy regularization and weighted functional analytic methods. With entropy regularization, we characterize a regular relaxed equilibrium through a fixed-point operator. By introducing two weight functions with distinct roles, one controlling the growth of rewards and values and the other defining the ambient weighted space, we construct a compact invariant set under a product topology and apply the Schauder-Tychonoff fixed-point theorem to establish existence of regularized equilibria. Importantly, the invariant set can be chosen uniformly for small entropy weight $λ\in(0,1]$. We then let $λ\to0+$ and show, through compactness, concentration of Gibbs policies, and uniform-integrability arguments, that a subsequential limit is a relaxed equilibrium of the original unregularized problem. We further study a policy iteration algorithm (PIA) for the entropy-regularized equilibrium problem. Under a weighted-discounting structure and sufficiently strong discounting, we establish exponential convergence and uniqueness of the regularized equilibrium in a suitable weighted Banach space. Combining the policy-iteration error with a quantitative soft-max approximation bound, we show that the iterated policies constitute weighted $\varepsilon$-equilibria for the original unregularized problem and derive an explicit regret estimate. A numerical example illustrating the convergence of PIA under strong discounting and a counterexample demonstrating its failure under weak discounting are also provided.