AI 中文总结
本文提出连续延迟记忆随机梯度下降,在二维景观上比普通SGD探索更广、收敛更精确,并设计无需求解HJB方程的连续时间强化学习结构,其最优性条件恢复吉布斯策略。
AI 中文摘要
类星体是宇宙中光度较高的天体,其亮度呈现随机变化,这些变化编码了驱动它们的超大质量黑洞的信息。从地面巡天数据的时间序列(即光变曲线)中建模这些变化,是一个统计挑战。本文回顾了历史上如何将随机微分方程(SDEs)与神经网络参数化相结合以克服这一挑战。我们提出了连续延迟记忆随机梯度下降(Continuous-Delayed-Memory Stochastic Gradient Descent),该方法依赖于离散迭代过程的过去状态。我们在一些二维景观上进行了模拟,通过调整超参数,观察到与普通SGD(Vanilla SGD)相比,该方法具有更广泛的探索范围和更精确的收敛行为。此外,我们提出了一种具有连续时间策略梯度的强化学习结构,用于探索性策略,而无需求解HJB偏微分方程(HJB PDE),并证明其最优性条件恢复了先前工作中的吉布斯策略(Gibbs policy)。
英文摘要
Quasars are luminous objects in the universe that exhibit stochastic brightness variations encoding information about the supermassive black holes powering them, and modeling these variations from ground-based survey data time series, known as light curves, is a statistical challenge. This paper reviews how stochastic differential equations (SDEs) have been adapted with neural network parameterizations to overcome this challenge in history. We create the Continuous-Delayed-Memory Stochastic Gradient Descent which depend on the past state of the discrete iteration process. We performed the simulation on some 2-dimensional landscape and observed some wider-exploration and more precise convergent behavior compared to Vanilla SGD by adjusting hyperparameters. Besides, we proposed a reinforcement learning structure with continuous time policy gradients for exploratory policies without solving HJB PDE, and we show that its optimality conditions recover the Gibbs policy of previous works.
CommentsKeywords: Stochastic process, Stochastic gradient descent, Continuous-Delayed-Memory Stochastic Gradient Descent, Stochastic Delay Differential Equation, Reinforcement Learning, Adjoint method