用于投资组合优化的科学物理信息强化学习
SciPhy Reinforcement Learning for Portfolio Optimization
AI总结:
研究为大型机构投资者用SciPhyRL构建动态投资组合优化框架,利用离线数据学习策略,核心是将优化转化为求解HJB方程并用PINN求解,控制变量转换后评估效果良好,能将信号质量转化为有效分配机制。
AI中文摘要:
本文为大型机构投资者引入了一个使用科学物理信息强化学习(SciPhyRL)的动态投资组合优化框架。该方法在包含明确累积成本的扩展状态空间中以连续时间制定,利用离线历史数据学习最优的、分布感知策略。核心创新是通过将其投影到观测轨迹上作为路径哈密顿 - 雅可比方程,将优化挑战简化为求解HJB方程,使用PINN在单次离线扫描中直接从数据求解,无需传统的值迭代或策略迭代。为使方法在实际短时间范围内有效,将控制变量从连续交易率转换为离散目标持有量。在14资产ETF领域使用设计的预言信号进行评估,学习到的吉布斯策略在样本外夏普比率上比静态和近视基线有显著提高。结果表明,所提出的框架成功地将已知信号质量转化为具有严格控制的波动性和周转率的强大、多期且成本感知分配机制。
英文摘要:
This paper introduces a dynamic portfolio optimization framework for large institutional investors using Scientific Physics-Informed Reinforcement Learning (SciPhyRL). Formulated in continuous time over an extended state space that includes explicit cumulative costs, the approach leverages offline historical data to learn optimal, distribution-aware strategies. A core innovation reduces the optimization challenge to solving an HJB equation by projecting it onto observed trajectories as a pathwise Hamilton-Jacobi equation. This is solved directly from data using PINN in a single offline sweep, eliminating the need for traditional value or policy iteration. To make the method effective at practical short horizons, the control variable is recast from a continuous trading rate to a discrete target holding. This ensures signal-implied positions are reached immediately, while execution costs are evaluated against a microstructure-grounded quadratic price impact model. Evaluated on a $14$-asset ETF universe using an engineered oracle signal, the learned Gibbs policy yields substantial out-of-sample Sharpe ratio improvements over static and myopic baselines. The results demonstrate that the proposed framework successfully translates known signal quality into a robust, multi-period, and cost-aware allocation mechanism with strictly controlled volatility and turnover.