发表机构
School of Electrical and Electronic Engineering, University College Dublin; School of Vehicle and Mobility, Tsinghua University; College of Artificial Intelligence, Tsinghua University(都柏林大学学院电气与电子工程学院; 清华大学车辆与运载学院; 清华大学人工智能学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究分布软策略迭代中策略评估步骤里的固定策略分布软贝尔曼算子,基于克拉默几何,制定CDF级算子并证明其收缩性质,获得唯一不动点,还通过共轭得到谱域表示,为DSPI式算法研究提供参考。
AI 中文摘要
分布软策略迭代(DSPI)为结合分布强化学习(DRL)与最大熵控制提供了重要框架,其策略评估步骤由作用于熵正则化回报的分布软贝尔曼算子主导。本文聚焦于基于累积分布函数(CDF)且具有\(L^2\)结构的克拉默几何,研究固定策略分布软贝尔曼算子在此度量下是否具有收缩性质及唯一不动点。直接在允许的CDF场域上,制定CDF级分布软贝尔曼算子,证明其为\(\sqrt{\gamma}\)收缩,并获得相应唯一不动点及收敛的迭代策略评估。CDF公式表明此有限克拉默域性质源于联合一步奖励熵转移的均匀一阶矩条件。通过共轭将评估问题转移到谱域,得到相同决策过程的等效希尔伯特空间表示。这些结果确定了与DSPI策略评估步骤相关的克拉默几何贝尔曼不动点,为研究DSPI式算法中的近似评论家、评估误差和评论家损失设计提供了参考点。
英文摘要
Distributional soft policy iteration (DSPI) provides an important framework for combining distributional reinforcement learning (DRL) with maximum-entropy control, in which the policy evaluation step is governed by a distributional soft Bellman operator acting on entropy-regularised returns. Theoretical analysis of such an evaluation step requires a probability metric under which Bellman updates can be controlled, typically by showing that the operator contracts the distance between any two candidate return-distribution estimates. In this paper, we focus on the Cramér geometry, a cumulative distribution function (CDF)-based metric with an $L^2$ structure, and study whether the fixed-policy distributional soft Bellman operator has this contraction property and hence a unique fixed point under this metric. Working directly on an admissible CDF field domain, we formulate the CDF-level distributional soft Bellman operator, prove that it is a $\sqrtγ$-contraction, and obtain the corresponding unique fixed point together with convergent iterative policy evaluation. The CDF formulation also shows that this finite-Cramér-domain property follows from a uniform first-moment condition on the combined one-step reward entropy shift, rather than from separate uniform boundedness assumptions on the reward and entropy terms. We then transport the same evaluation problem to the spectral domain by conjugation, obtaining an equivalent Hilbert-space representation of the same decision process. Taken together, these results identify the Cramér-geometric Bellman fixed point associated with the policy-evaluation step of DSPI, providing a reference point for studying approximate critics, evaluation error, and critic-loss design in DSPI-style algorithms.