AI 中文总结
研究提出CEFOL算法解决含递归效用的离散时间动态规划问题,通过引入神经网络表示确定性等价,利用一阶条件构建残差学习相关函数,能处理多种约束,应用于多个问题,样本外诊断和残差表现良好,对含期望效用问题也适用。
AI 中文摘要
本文提出确定性等价一阶学习(CEFOL)算法,用于解决含递归效用的离散时间动态规划问题。含递归效用的动态规划具有挑战性,因贝尔曼方程和一阶最优性条件中出现非线性确定性等价且难评估。CEFOL通过引入单独神经网络表示确定性等价,利用贝尔曼方程和特定模型一阶最优性条件。它还通过特定模型一阶条件构建残差来学习价值函数、策略函数和拉格朗日乘数。通过使用一阶和KKT残差学习策略,CEFOL可直接处理控制上的一般等式和不等式约束。将该算法应用于风险敏感和爱泼斯坦 - 津消费储蓄问题等,样本外贝尔曼诊断和特定模型最优性残差在相关状态区域通常为1.0e - 4到1.0e - 3,学习的价值和策略函数与VFI基准紧密匹配。CEFOL算法对含期望效用的动态规划问题也有效,因为期望效用是递归效用的特殊情况。
英文摘要
This paper proposes the certainty-equivalent first-order learning (CEFOL) algorithm, a deep learning algorithm for solving discrete-time dynamic programming problems with recursive utility. Dynamic programming with recursive utility is challenging because nonlinear certainty equivalent appears in the Bellman equation and the first-order optimality conditions but is difficult to evaluate. By introducing a separate neural network to represent the certainty equivalent, CEFOL enables the exploitation of the Bellman and model-specific first-order optimality conditions. In addition to certainty equivalent, CEFOL also uses neural networks to learn the value functions, policy functions, and Lagrange multipliers by using model-specific first-order conditions to construct residuals for minimization. By using first-order and KKT residuals to learn the policy, CEFOL directly accommodates general equality and inequality constraints on the controls, including occasionally binding constraints, without requiring penalty functions or problem-specific reformulations. We apply the algorithm to risk-sensitive and Epstein--Zin consumption-saving problems, a small-noise robust-control problem, and a DSGE model with recursive preferences and stochastic volatility. Across these applications, out-of-sample Bellman diagnostics and model-specific optimality residuals, including Euler or first-order residuals where applicable, are generally of order 1.0e-4 to 1.0e-3 over the relevant state regions, with larger values mainly near binding constraints, and the learned value and policy functions closely match VFI benchmarks when available. The CEFOL algorithm also works for dynamic programming problems with expected utility, as expected utility is a special case of recursive utility.
Comments86 pages, 44 figures