arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

当正确奖励不足时:在解析求解的经纪商-交易商博弈中诊断与引导PPO

When a Correct Reward Is Not Enough: Diagnosing and Guiding PPO in an Analytically Solved Broker-Trader Game

Siu Tung Wong, Carlo Campajola

arXiv 2610.03598首次发表:更新:

发表机构

University College London; UZH Blockchain Center(伦敦大学学院; 苏黎世大学区块链中心)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究在解析求解的经纪商-交易商博弈中诊断PPO,发现其critic排序能力不足,并利用解析策略作为基准和初始策略,通过微调实现成本变化后的有效适应。

AI 中文摘要

强化学习(RL)越来越多地用于金融最优控制问题,尤其是在复杂动态使得解析策略难以获得的情况下。金融数学文献提供了许多已求解模型,其方程和控制可用于评估和引导学习;我们探讨RL能否利用这些结果。我们将近端策略优化(PPO)智能体置于一个解析求解的连续时间经纪商-交易商博弈中。PPO替代经纪商,在与知情交易者和随机不知情订单流交互时选择其交易速度。我们从经纪商的连续时间收益中推导出有限步奖励,并通过网格细化和精确的一步恒等式验证其离散实现。在零不知情流的情况下,验证选择的PPO-FFNN接近参考动作。在随机不知情流下,测试的PPO-FFNN和PPO-LSTM仍不准确,尽管监督学习确认其actor能够表示该动作。蒙特卡洛诊断显示,其critic不能可靠地对邻近动作进行排序;基于势能的奖励塑形也未提供可靠改进。在部分信息下,基于经纪商可观察历史的因果确定性等价控制器仍接近参考,而PPO具有较大误差和较低收益。最后,我们冻结解析策略,并在执行成本变化后训练PPO进行调整。将成本减半产生了可重复的改进,缩小了与变化成本参考差距的2.22%。因此,解析解既提供了诊断RL的基准,也为适应提供了有用的起始策略。

英文摘要

Reinforcement learning (RL) is increasingly used for financial optimal-control problems when complex dynamics make analytical strategies difficult to obtain. There are financial mathematics literactures which provides many solved models whose equations and controls could evaluate and guide learning; we ask whether RL can exploit these results. We place a proximal policy optimisation (PPO) agent in an analytically solved continuous-time broker--trader game. PPO replaces the broker and chooses its trading speed while interacting with an informed trader and stochastic uninformed order flow. We derive a finite-step reward from the broker's continuous-time payoff and verify its discrete implementation through grid refinement and an exact one-step identity. With zero uninformed flow, a validation-selected PPO--FFNN approaches the reference action. With stochastic uninformed flow, the tested PPO--FFNN and PPO--LSTM remain inaccurate, although supervised learning confirms that their actors can represent the action. Monte Carlo diagnostics show that their critics do not reliably rank nearby actions; potential-based reward shaping also gives no reliable improvement. Under partial information, a causal certainty-equivalent controller based on the broker's observable history remains close to the reference, while PPO has larger errors and lower payoffs. Finally, we freeze the analytical policy and train PPO to adjust it after the execution cost changes. Halving the cost yields a repeatable improvement that closes \(2.22\%\) of the gap to the changed-cost reference. The analytical solution therefore provides both a benchmark for diagnosing RL and a useful starting policy for adaptation.

Comments8 pages; accepted for publication at ICAIF 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑