arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.22785cs.LG

基于Hawkes订单流与价格冲击的对抗强化学习稳健做市策略

Robust Market Making with Hawkes Order Flow and Price Impact via Adversarial Reinforcement Learning

  • North China Institute of Computer System Engineering(华北计算机系统工程研究所)
  • University of Science and Technology of China(中国科学技术大学)

机构由 AI 辅助整理,请以论文原文为准。

Hao Yang, Zhenguo Xu

AI总结:

针对做市策略面临模型不确定性和制度风险,提出结合Hawkes订单流、价格冲击与LSTM的对抗强化学习框架,通过博弈论分析和实验验证,在复杂微观结构环境中显著提升收益左尾稳健性。

AI中文摘要:

在真实限价订单簿市场中的做市策略面临显著的模型不确定性和制度转换风险。现有的对抗强化学习方法通过将Avellaneda-Stoikov做市问题构建为做市商与环境对手之间的零和博弈来提高稳健性。然而,这些方法通常依赖于泊松订单到达,并忽略交易引发的价格冲击,限制了其捕捉重要高频市场微观结构效应(如聚集订单流、自激发和交易后价格反馈)的能力。我们将做市对抗强化学习扩展到更复杂的环境,该环境包含Hawkes自激发订单到达和交易引发的价格冲击。为缓解扩展制度空间引入的更强非平稳性,我们引入LSTM模块,显式建模近期观测的时间结构。我们进一步通过博弈论分析和数值实验刻画所提出框架的均衡性质,并引入聚焦于收益分布左尾改进的稳健性评估协议。跨多种市场制度的实验结果表明,所提方法在大多数复杂微观结构环境中实现了左尾性能的改进。特别地,在具有强Hawkes激发和低至中等价格冲击的制度下,改进尤为显著。Bootstrap检验未提供证据表明这些改进是通过更强的终端方向性库存偏差获得的。这些结果表明,将对抗训练与时间状态表示相结合,可以提高基于强化学习的做市策略在订单流自激发、价格冲击和制度不确定性下的稳健性。

英文摘要:

Market-making strategies in real limit order book markets face substantial model uncertainty and regime-shift risk. Existing adversarial reinforcement learning approaches improve robustness by formulating the Avellaneda--Stoikov market-making problem as a zero-sum game between a market maker and an environmental adversary. However, these approaches typically rely on Poisson order arrivals and neglect trade-induced price impact, limiting their ability to capture important high-frequency market microstructure effects such as clustered order flow, self-excitation, and post-trade price feedback. We extend adversarial reinforcement learning for market making to a more complex environment with Hawkes self-exciting order arrivals and trade-induced price impact. To mitigate the increased non-stationarity introduced by the expanded regime space, we incorporate an LSTM module that explicitly models the temporal structure of recent observations. We further characterize the equilibrium properties of the proposed framework through both game-theoretic analysis and numerical experiments, and introduce a robustness evaluation protocol focused on improvements in the left tail of the return distribution. Experimental results across a range of market regimes show that the proposed method achieves improved left-tail performance in most complex microstructure environments. In particular, the gains are pronounced in regimes with strong Hawkes excitation and low-to-moderate price impact. Bootstrap tests provide no evidence that these improvements are obtained through a stronger terminal directional inventory bias. These results suggest that combining adversarial training with temporal state representation can improve the robustness of reinforcement-learning-based market-making strategies under order-flow self-excitation, price impact, and regime uncertainty.

补充信息

↑