何时不作为是错误?面向PPO交易策略的延续感知审计
When Is Inaction a Mistake? Continuation-Aware Auditing of PPO Trading Policies
浏览论文内容
中文总结 AI 辅助
针对冻结的PPO交易策略,提出四阶段延续感知审计方法,通过信息匹配与延续分析区分动作更改与策略替换,在模拟和BTCUSDT回放中显著提升收益并降低换手成本。
中文摘要 AI 辅助
当学习到的策略选择不作为(不执行)时,最优参考可能建议进行交易,但该建议取决于信息和未来决策。我们提出了一种针对冻结的近端策略优化(PPO)策略的四阶段审计方法,无需重新训练。该方法检查部署占用率,匹配当前信息,在现有延续策略下测试孤立偏差,并评估基于观测的替代方案的重复部署。在受控的线性高斯模拟中,信息匹配解释了部分分歧,而延续策略改变了其解释。在单位观测噪声下,现有延续策略逆转了99.3%的预计困难错失优势质量;重复的预计规则部署改善了全部50个策略。这些比较区分了孤立动作更改与策略替换。历史比特币/泰达币(BTCUSDT)回放将这种部署视角应用于一个使用2024年数据选择并冻结至2025年的人工指定干预。每日净收益提高了135.03个基点,在50个策略中有46个获得收益,主要归因于更低的换手成本。该审计阐明了预言机标记的不作为(不执行)对部署决策制定的意义。
英文摘要
An optimal reference may recommend trading when a learned policy chooses inaction, but the recommendation depends on information and future decisions. We introduce a four-stage audit for frozen proximal policy optimization policies without retraining. It examines deployment occupancy, matches current information, tests isolated deviations under incumbent continuation, and evaluates repeated deployment of observation-based alternatives. In controlled linear-Gaussian simulations, information matching explains part of the disagreement, while continuation changes its interpretation. At unit observation noise, incumbent continuation reverses 99.3% of projected-hard missed-advantage mass; repeated projected-rule deployment improves all 50 policies. These comparisons distinguish isolated action changes from policy replacement. Historical Bitcoin/Tether (BTCUSDT) replay applies this deployment perspective to a hand-specified intervention selected using 2024 data and frozen for 2025. Daily net reward improves by 135.03 basis points, with gains in 46 of 50 policies, primarily through lower turnover costs. The audit clarifies what oracle-flagged inaction implies for deployed decision making.