面向CVaR风险感知Q学习的自适应有限预算训练
Adaptive Finite-Budget Training for CVaR Risk-Aware Q-Learning
- The University of Hong Kong(香港大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本文针对CVaR风险感知Q学习提出自适应有限预算训练控制器,通过六个协同机制优化训练流程,在比特币交易任务中大幅降低贝尔曼残差、提升策略风险调整后表现。
AI中文摘要:
风险感知Q学习(RaQL)为动态风险目标提供了一种无模型的双时间尺度估计器,但其有限预算表现仍不稳定:固定的内环超参数会产生不稳定的价值估计、持续的贝尔曼残差以及低效的样本复用。本文针对条件风险价值(CVaR)RaQL提出了一种自适应训练控制器,并在每日比特币交易任务上进行了评估。该控制器保留了原始CVaR估计器和贝尔曼不动点,而是通过六个协同机制重新设计了训练流程:逐单元内步长调整、与外环速率匹配的衰减同步、针对类VaR内变量的短期早期校正、先覆盖后贪心的样本分配规则、成熟内估计的渐进后缀聚合,以及基于在线可观测量的关键尺度数据驱动校准。在20个随机种子和856000个内转移样本的条件下,该控制器相较于固定参数基线,将平均经验CVaR贝尔曼残差降低了约85%(MeanBEQ从1.2202降至0.1854;MeanBEV从1.1624降至0.0535),并在不同CVaR水平、折扣因子和训练预算下保持稳定。在时间划分的样本外测试集上,学习到的策略在扣除交易成本后获得了0.9281的夏普比率和6.46%的最大回撤。尽管买入并持有策略的累计收益更高(35.43% vs. 23.61%),但自适应策略实现了低得多的波动率(9.57% vs. 47.93%)、回撤和CVaR损失。这些结果表明,仅应用于训练流程而不改变风险目标的自适应有限预算训练设计,可显著提升风险感知Q学习在金融应用中的可靠性和风险调整后表现。
英文摘要:
Risk-aware Q-learning (RaQL) provides a model-free, two-timescale estimator for dynamic risk objectives, but its finite-budget behavior remains fragile: fixed inner-loop hyperparameters can produce unstable value estimates, persistent Bellman residuals, and inefficient sample reuse. This paper proposes an adaptive training controller for Conditional Value-at-Risk (CVaR) RaQL and evaluates it on a daily Bitcoin trading task. The controller preserves the original CVaR estimator and Bellman fixed point; instead, it redesigns the training procedure through six coordinated mechanisms: per-cell inner-step sizing, outer-rate-matched decay synchronization, a short early correction for the VaR-like inner variable, a coverage-first-then-greedy sample allocation rule, progressive suffix aggregation of mature inner estimates, and data-driven calibration of key scales from online-observable quantities. Across 20 random seeds and 856,000 inner-transition samples, the controller reduces the mean empirical CVaR Bellman residual by approximately 85% relative to the fixed-parameter baseline (MeanBEQ: 1.2202 to 0.1854; MeanBEV: 1.1624 to 0.0535) and maintains stability across CVaR levels, discount factors, and training budgets. On the chronological out-of-sample test set, the learned policy attains a Sharpe ratio of 0.9281 with a maximum drawdown of 6.46% after transaction costs. Although buy-and-hold yields a higher cumulative return (35.43% vs. 23.61%), the adaptive policy achieves far lower volatility (9.57% vs. 47.93%), drawdown, and CVaR loss. These results demonstrate that adaptive finite-budget training design, applied solely to the training procedure without altering the risk objective, can materially improve the reliability and risk-adjusted performance of risk-aware Q-learning in financial applications.