arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

股票策略回测:运气还是优势?MinervaScore作为统计稳健性评级

Equity Strategy Backtesting: Luck or Edge? The MinervaScore as a Statistical Robustness Grade

Maria Laura Santoni, Vincent Jouanne, Matthew L. Scullin

arXiv 2608.23808首次发表:更新:

AI 中文总结

针对交易策略回测易受运气影响的问题,提出MinervaScore稳健性评级,结合多指标与制度稳定性诊断,能区分真实信号与幸运结果,可作为可审计的验证报告层。

AI 中文摘要

交易策略的回测通常是在多次参数尝试后才被选中的,因此出色的历史结果可能反映的是搜索运气而非持续信号。收益率、夏普比率和最大回撤等标准指标未记录尝试过的候选策略数量、所选规则是否通过样本外验证,以及可用历史数据是否足够支撑结果。本文介绍了MinervaScore,一种用于交易策略的后选择稳健性评级。该评级结合了四个成熟的验证指标:去偏夏普比率(Deflated Sharpe Ratio)、回测过拟合概率(Probability of Backtest Overfitting)、 Superior Predictive Ability( superior预测能力)和最小跟踪记录长度,以及一个 regime-stability(制度稳定性)诊断指标。这些组件被转换为相对于其可接受阈值的有符号边际,汇总为原始分数后映射到0-100的显示值,该显示值与一个二元稳健性印章绑定:仅当所有五个关卡都通过时,才显示80分及以上的评级。校准使用了359,062条生产级回测记录。该评级旨在对统计支撑进行排名,而非估计未来盈利的概率。在具有已知真实值的合成市场中,MinervaScore能将真实信号与幸运的回测结果区分开,在主要难度下的AUROC为0.989。它相对于GT-Score代理指标和关卡通过基线的提升幅度不大,且与仅去偏夏普比率的校正基线接近。在对未见过的真实市场数据进行的预注册测试中,该评级在优势有限的总体中未显示出显著的正向关系(斯皮尔曼相关系数ρ_s=0.013,单侧置换检验p=0.40)。因此,我们将MinervaScore作为可审计的验证和报告层,而非实际市场可预测性的证据。

英文摘要

Backtests of trading strategies are often selected after many parameter trials. A strong historical result can therefore reflect search luck rather than a persistent signal. Standard summaries such as return, Sharpe ratio, and drawdown do not record how many candidates were tried, whether the selected rule survives out-of-sample validation, or whether the available history is long enough to support the result. This paper describes the MinervaScore, a post-selection robustness grade for trading strategies. The score combines four established validation quantities: Deflated Sharpe Ratio, Probability of Backtest Overfitting, Superior Predictive Ability, and Minimum Track Record Length, with a regime-stability diagnostic. These components are converted into signed margins from their admissibility thresholds, aggregated into a raw score, and then mapped to a 0-100 display. The display is tied to a binary Robustness Seal: scores of 80 or higher are shown only when all five gates pass. The calibration uses 359,062 production backtest records. The score is intended to rank statistical support, not to estimate the probability of future profit. In synthetic markets with known ground truth, the MinervaScore separates true signal from lucky backtest outcomes, with an AUROC of 0.989 at the headline difficulty. Its improvement over the GT-Score proxy and the gates-passed baseline is modest, and it remains close to the corrected DSR-alone baseline. In a pre-registered test on unseen real-market data, the score showed no significant forward relationship in a population with limited surviving edge (Spearman rho_s = 0.013, one-sided permutation p = 0.40). We therefore present the MinervaScore as an auditable validation and reporting layer, rather than as evidence of demonstrated real-market predictability.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑