发表机构
Sungkyunkwan University; Seoul National University of Science and Technology; Chonnam National University(成均馆大学; 首尔科学技术大学; 全南国立大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究针对带可预测收益和凸约束的连续时间投资组合选择问题,提出自洽伴随策略迭代方法,经理论证明和数值实验验证其收敛性与性能优势。
AI 中文摘要
我们开发了基于模拟的策略迭代方法,用于处理具有可预测收益和凸约束的连续时间投资组合选择问题。每一轮外循环步骤会在部署后重新评估固定隐式OL-BPTT伴随,并求解受约束更新。移位伴随抵消通过策略改进残差控制伴随-HJB哈密顿量梯度的差异。对于CRRA投资组合,精确HJB策略迭代可识别最优约简价值因子,而总体OL-BPTT迭代在占据测度相对误差条件下全局收敛。经定理匹配的检验得到所需0.75阈值对应的最大95%上限端点为0.074。在三因子、五十资产的设计中,当前策略重评估在两种评估规则下均优于匹配的池化优化。
英文摘要
We develop simulation-based policy iteration for continuous-time portfolio choice with predictable returns and convex constraints. Each outer step re-evaluates a fixed-latent open-loop backpropagation-through-time (OL-BPTT) adjoint after deployment and solves the constrained update. Shifted-adjoint cancellation controls the adjoint--HJB Hamiltonian-gradient discrepancy by the policy-improvement residual. For CRRA portfolios, exact HJB policy iteration identifies the optimal reduced value factor, while population OL-BPTT iteration converges globally when the adjoint update is directionally improving and approximate stationarity is asymptotically HJB-compatible. A theorem-matched occupation audit yields maximal $95\%$ upper endpoints of $0.066$ for the primitive directional ratio and $0.074$ for a stronger norm-relative ratio, both against the half-step threshold $0.75$. In the high-precision $50$--$50$ occupancy/broad-anchor design of a three-factor, fifty-asset benchmark, current-policy re-evaluation outperforms matched pooled refinement under the on-policy and broad evaluation laws.