面向受约束动态投资组合选择的可扩展庞特里亚金引导伴随到控制恢复
Scalable Pontryagin-Guided Adjoint-to-Control Recovery for Constrained Dynamic Portfolio Choice
浏览论文内容
中文总结 AI 辅助
该研究针对带平滑逐点约束的连续时间投资组合选择,开发了庞特里亚金引导的可扩展伴随到控制框架,通过OL-BPTT等方法实现了高精度恢复,适用于含100个风险资产的投资组合。
中文摘要 AI 辅助
我们开发了一种用于带平滑逐点约束的连续时间投资组合选择的可扩展伴随到控制框架。可行的直接策略优化(DPO)策略提供滚动输出;训练后,固定潜伏开环反向传播通过时间(OL-BPTT)产生一阶和二阶路径敏感性,其条件投影生成适配的伴随输入。嵌套对偶公共随机数回归估计移位财富行鞅输入,部署时通过二次仿射块的精确二次规划(QP)或其他情况下的对数障碍求解局部广义哈密顿问题。我们证明了保留正交投影残差的OL-BPTT-庞特里亚金极大值原理(PMP)对应关系、局部障碍-卡尔什-库恩-塔克(KKT)近似,以及局部二次增长下的性能到伴随桥梁。在n=100的受约束默顿基准中,学习到的一阶伴随具有0.46%的平均相对误差;在解析策略下,一阶伴随、财富曲率和布朗系数的归一化均方根误差(nRMSE)分别为0.031%、0.035%和0.326%。在仅终端可预测回报的固定相对风险厌恶(CRRA)基准中,财富齐次性给出$P^{X\bullet,*}=D^2_{X\bullet}V$和$\boldsymbol{\beta}^{X,*}=0$。在512×16的主要投影预算下,完整估计移位解码器的策略均方根误差(RMSE)低于$8.5\times10^{-3}$,而特定基准的零移位神谕低于$2\times10^{-4}$。在约束、因子、障碍和切换区域的审计中,该恢复方法大幅降低了局部KKT残差,且对多达100个风险资产的投资组合仍具实用性。
英文摘要
We study continuous-time multi-asset portfolio choice and consumption under smooth pointwise constraints, including state-dependent feasible sets. The method separates dynamic information acquisition from local constrained recovery. A pointwise-feasible neural actor generates reference rollouts; after training, its realized latent outputs are frozen and first- and second-order adjoints are harvested from a fixed-latent open-loop backpropagation-through-time graph. Feedback therefore generates the reference trajectory without restricting the adjoint formulation to Markov controls. Conditional on the harvested adjoints, deployment solves a local generalized Pontryagin-Hamiltonian problem: quadratic-affine portfolio blocks are recovered exactly by a quadratic program, while a log barrier approximates more general regular KKT branches. We establish local chart representations, an OL-BPTT-to-adjoint correspondence retaining orthogonal martingale residuals, and an end-to-end bound from reference value loss and numerical errors to recovered-policy and local QP-gap errors. Analytical constant- and predictable-opportunity benchmarks validate the adjoints. Common-input experiments show that recovery reduces residual PMP/KKT error left by finite-budget direct policy optimization, including under a state-dependent consumption cap and with up to 100 risky assets. Scalability concerns the constrained action block rather than dimension-free state-space complexity.
发表机构
- Graduate School of Data Science, Chonnam National University(全南大学数据科学研究生院)
- Department of Mathematics, Sungkyunkwan University(成均馆大学数学系)
- Department of Financial Engineering, Ajou University(亚洲大学金融工程系)
- Department of Fintech, SKK Business School, Sungkyunkwan University(成均馆大学SKK商学院金融科技系)
机构由 AI 辅助整理,请以论文原文为准。