arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

基于广义卡尔曼滤波器的时间差分强化学习

Generalized Kalman filter based temporal difference reinforcement learning

Vasos Arnaoutis, Eric Lutters, Bojana Rosić

arXiv 2607.20010首次发表:更新:

发表机构

University of Twente; University of Vienna(特温特大学; 维也纳大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究基于条件期望理论的广义TD强化学习框架,将价值和Q值函数视为不确定量,用随机推理估计,能递归估计价值函数条件期望及二阶概率矩,通过离散化随机问题验证,可扩展经典卡尔曼时间差分学习到更广泛随机系统。

AI 中文摘要

本文提出了一种基于条件期望理论的广义时间差分(TD)强化学习框架。将价值函数和动作价值(Q值)函数视为不确定量,其估计被公式化为一个随机推理问题。与依赖线性高斯假设的经典基于卡尔曼的时间差分学习不同,该公式直接从条件期望框架导出,自然扩展到非线性模型和非高斯概率分布。该方法不仅递归估计价值函数的条件期望,还估计其二阶概率矩,从而在整个学习过程中量化与学习到的价值函数相关的不确定性。为获得计算上易于处理的算法,使用多项式混沌展开或基于集合的近似对随机问题进行离散化,提供基础随机变量的有效表示。该框架在两个最优控制问题上得到验证:线性质量-弹簧-阻尼器系统和封闭腔内的非线性热传导问题。数值例子说明了该方法准确估计价值函数及其相关不确定性的能力,同时将经典基于卡尔曼的时间差分学习扩展到更广泛的随机系统类别。

英文摘要

In this paper, we present a generalized temporal-difference (TD) reinforcement learning framework based on the theory of conditional expectations. The value and action-value (Q-value) functions are treated as uncertain quantities, and their estimation is formulated as a stochastic inference problem. Unlike classical Kalman-based temporal-difference learning, which relies on linear-Gaussian assumptions, the proposed formulation is derived directly from the conditional expectation framework and naturally extends to nonlinear models and non-Gaussian probability distributions. The proposed method recursively estimates not only the conditional expectation of the value function but also its second probabilistic moment, thereby quantifying the uncertainty associated with the learned value function throughout the learning process. To obtain a computationally tractable algorithm, the stochastic problem is discretized using either polynomial chaos expansions or ensemble-based approximations, providing efficient representations of the underlying random variables. The proposed framework is demonstrated on two optimal control problems: a linear mass--spring--damper system and a nonlinear heat conduction problem in a closed cavity. The numerical examples illustrate the capability of the proposed method to accurately estimate both the value function and its associated uncertainty, while extending classical Kalman-based temporal-difference learning to a broader class of stochastic systems.

Comments39 pages, 18 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑