arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

基于机制切换订单流的稳健做市深度强化学习

Deep Learning of Robust Market Making under Regime-Switching Order Flow

Felipe Moret, Fabrizio Lillo

arXiv 2609.11614首次发表:更新:

发表机构

Scuola Normale Superiore di Pisa(比萨高等师范学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出基于Rainbow风格C51的深度强化学习做市商(RLMM),通过增加贝叶斯变点滤波器和队列调整不平衡信号,在机制切换订单流下优于GLFT并恢复盈利。

AI 中文摘要

基于随机控制的经典做市策略,如Avellaneda-Stoikov模型及Guéant-Lehalle-Fernandez-Tapia(GLFT)扩展,提供了闭式报价规则,但其假设在现实微观结构时间尺度上不成立。其中一个假设是订单流是平稳的,而经验证据表明存在机制(regimes),可能与元订单的算法执行相关。在这种情况下,现有方法会产生负收益。本文开发了一种深度强化学习做市商(RLMM)——一种Rainbow风格的分位数DQN(C51),在零智能限价订单簿中进行了校准和测试。我们发现,在平稳环境下,RLMM在观测到的整个风险-收益前沿上优于GLFT。RLMM对订单流不对称性的鲁棒性优于GLFT,但与任何平稳训练的策略一样,在持续方向性失衡下,仍会因库存饱和而遭受大幅回撤。通过向RLMM的状态增加两个辅助信号——方向性流量偏差的贝叶斯在线变点滤波器和队列调整的报价暴露不平衡——可恢复盈利能力。最后,一个场景-老虎机步骤,对低回报机制场景进行重新加权,进一步提高了在随机持续性和相关方向压力下的性能。

英文摘要

Classical market-making strategies based on stochastic control, such as the Avellaneda-Stoikov and the Guéant-Lehalle-Fernandez-Tapia (GLFT) extension, provide closed-form quoting rules, but rest on assumptions that break down at realistic microstructure timescales. One of them is that order flow is stationary, while empirical evidence points to the existence of regimes, possibly associated with algorithmic execution of metaorders. In this case, existing methods provide negative PnL. In this paper, we develop a deep reinforcement-learning market maker (RLMM) - a Rainbow-style distributional DQN (C51) which is calibrated and tested in a zero-intelligence limit order book. We find that, in the stationary setting, RLMM outperforms GLFT across the entire observed risk-return frontier. The RLMM is more robust to flow asymmetry than GLFT, but, like any stationarily trained strategy, it still suffers large drawdowns from inventory saturation under persistent directional imbalance. Augmenting the state of RLMM with two auxiliary signals - a Bayesian online change-point filter over the directional flow bias and a queue-adjusted quote-exposure imbalance -restores profitability. A final scenario-bandit step that reweights low-return regime scenarios further improves performance under random-persistence and correlated-direction stress.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑