基于模拟的投资组合选择神经策略:架构、训练与可解释性
Simulation-Based Neural Policies for Portfolio Choice: Architecture, Training, and Interpretability
浏览论文内容
中文总结 AI 辅助
该研究针对投资组合选择问题,设计并比较四种神经策略架构,结合模拟优化与多维度评估,解决维度诅咒问题,为低维场景提供可诊断的架构与训练方案。
中文摘要 AI 辅助
许多经济决策问题,如生命周期消费-储蓄和动态投资组合选择,是具有连续状态和动作的有限时间随机控制问题。当状态为低维时,这些问题可通过网格上的动态规划求解。网格计算成本随状态维度呈指数增长,即所谓的“维度诅咒”,这促使人们用通过模拟直接优化的神经策略替代价值函数网格。这类策略通常在其适用的高维场景中研究,而这些场景恰好不存在参考解,因此无法分离和诊断任何单一架构或训练选择的贡献。为此,我们将架构和求解方法均作为研究对象,考虑一个具有足够低维归一化状态空间的生命周期问题,该问题可获得精确的动态规划解,用于评估。我们比较四种架构:最简单的是单个带时间条件的网络;然后考虑两个在制度切换处连接的网络,接着是每个日期对应一个网络,在冻结下游策略的情况下反向训练;最后评估按日期架构的约束变体。将策略按时间解耦,使每个日期具有简短且适定的目标,我们将其与归一化梯度幅度的方向主导优化相结合。导致相似实现效用目标的架构,在是否尊重潜在问题的经济学原理方面可能存在差异,因此我们基于福利、无解的贝尔曼残差、形状约束以及生成的策略函数,对每种设计进行联合评估。
英文摘要
Many economic decision problems, lifecycle consumption-saving and dynamic portfolio choice, are finite-horizon stochastic control problems with continuous states and actions. When the state is low-dimensional these problems are solved by dynamic programming on a grid. The grid cost grows exponentially in the state dimension, known as the curse of dimensionality, which motivates replacing the value-function grid with a neural policy optimized directly through simulation. Such policies are usually studied in the high-dimensional settings that motivate them, precisely where no reference solution exists. So the contribution of any single architectural or training choice cannot be isolated and diagnosed. We therefore take a step back and treat both the architecture and the solution method as the objects of study. To this end, we consider a lifecycle problem with a sufficiently low-dimensional normalized state space to admit an accurate dynamic programming solution, which is used for evaluation. We compare four architectures. The simplest consists of a single time-conditioned network. We then consider two networks concatenated across the regime switch, followed by one network per date trained backward against frozen downstream policies. Finally, we evaluate a constrained variant of the per-date architecture. Decoupling the policy across time gives each date a short, well-posed objective, which we pair with direction-dominant optimization that normalizes away gradient magnitude. Architectures that lead to similar realized utility objective can nevertheless differ in whether they respect the underlying problem's economics. We therefore evaluate each design jointly based on welfare, a solution-free Bellman residual, shape restrictions, and the resulting policy functions.