发表机构
College of Engineering and Computer Science, VinUniversity(VinUniversity工程与计算机科学学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
提出PPO-HRAP,结合近端策略优化与制度先验的混合策略,在SPY测试中显著降低回撤并提升风险调整收益,优于买入并持有。
AI 中文摘要
用于交易的强化学习常常难以在上涨参与和回撤控制之间取得平衡。仅关注利润的策略可能会在上涨趋势的资产上退化为被动做多敞口,而激进的风险惩罚奖励在波动期可能变得过于防御。本文提出了PPO-HRAP,一种混合制度感知策略,将近端策略优化与可解释的制度先验相结合。智能体同时观察市场特征和投资组合状态变量,接收由投资组合对数收益、基于VIX的回撤增加惩罚、目标敞口偏差和换手成本组成的奖励,并执行PPO演员输出与制度推导的目标敞口之间的混合动作。在2020-2022年SPY的保留测试窗口上,PPO-HRAP实现了27.62%的总收益、8.48%的年化收益、0.6447的夏普比率、0.8588的索提诺比率和0.4592的卡尔马比率,同时将最大回撤从买入并持有的34.10%降低至18.47%。在五个SPY随机种子下,PPO-HRAP保持稳定,平均总收益为$0.2725 \u00b1 0.0109$,平均夏普比率为$0.6219 \u00b1 0.0565$。在QQQ和DIA上的单次跨资产测试进一步表明,所提方法在所有三个报告的资产上总收益和夏普比率均排名第一。这些结果表明,将学习到的动作与波动率感知的制度先验相结合是改善风险调整后交易行为的一种实用方法,尽管当前策略仍存在高换手率,且超出SPY的跨资产稳健性仍仅限于单次运行证据。
英文摘要
Reinforcement learning for trading often struggles to balance upside participation with drawdown control. Profit-only policies can collapse toward passive long exposure on upward-drifting assets, while aggressively risk-penalized rewards can become too defensive during volatile periods. This paper proposes PPO-HRAP, a hybrid regime-aware policy that combines Proximal Policy Optimization with an interpretable regime prior. The agent observes both market features and portfolio-state variables, receives a reward combining portfolio log return, VIX-conditioned drawdown-increase penalty, target-exposure deviation, and turnover cost, and executes a blended action between the PPO actor output and a regime-derived target exposure. On the held-out 2020-2022 SPY test window, PPO-HRAP achieves 27.62% total return, 8.48% annualized return, 0.6447 Sharpe ratio, 0.8588 Sortino ratio, and 0.4592 Calmar ratio, while reducing maximum drawdown from 34.10% for Buy and Hold to 18.47%. Across five SPY seeds, PPO-HRAP remains stable with mean total return $0.2725 \pm 0.0109$ and mean Sharpe ratio $0.6219 \pm 0.0565$. Single-run cross-asset tests on QQQ and DIA further show that the proposed method ranks first on total return and Sharpe ratio for all three reported assets. These results suggest that blending learned actions with a volatility-aware regime prior is a practical way to improve risk-adjusted trading behavior, although the current policy still incurs high turnover and cross-asset robustness beyond SPY remains limited to single-run evidence.
Comments8 pages, 6 figures, 8 tables. Code: https://github.com/chikien07012006/PPO_Regime-Aware-Trading