arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.13825cs.AIq-fin.TR

ViperQ:基于拍卖市场理论的订单流模式识别用于强化学习交易

ViperQ: Order Flow Pattern Recognition via Auction Market Theory for Reinforcement Learning Trading

Asser Moustafa, Rares-Mihail Neagu, Jugal Kalita

首次发表
浏览论文内容

中文总结 AI 辅助

ViperQ利用拍卖市场理论特征构建强化学习交易状态表示,在真实机构数据上实现TSLA和NVDA的高回报,验证了微观结构模式在顺序决策中的有效性。

中文摘要 AI 辅助

学术文献中发表的强化学习交易系统绝大多数依赖于价格聚合状态表示(OHLCV柱)或限价订单簿深度特征,这使得从业者文献中的微观结构模式理论,即拍卖市场理论和市场轮廓,缺乏经过同行评审的计算实现。我们提出了ViperQ,一个强化学习系统,其状态表示明确基于拍卖市场理论原语构建:成交量控制点、价值区域位置、低成交量节点标志、累计成交量增量背离以及磁带速度特征,组装成一个20维Z归一化向量。两个近端策略优化智能体使用基于前景理论的不对称奖励函数进行训练,该函数以与卡尼曼和特沃斯基损失厌恶系数一致的幅度惩罚亏损持仓。在机构逐笔交易数据的留出十二个月分区上评估,这些数据是智能体从未见过的,ViperQ在零杠杆下实现了TSLA上+163.6%的投资回报率(-27.5%最大回撤,27,019笔交易)和NVDA上+116.5%的投资回报率(-47.8%最大回撤,12,892笔交易)。结果确立了拍卖市场理论特征作为金融时间序列上顺序决策的可处理结构化输入模态,并激励了微观结构感知策略学习的进一步工作。

英文摘要

Reinforcement learning trading systems published in the academic literature overwhelmingly rely on price-aggregate state representations (OHLCV bars) or limit-order-book depth features, leaving microstructure pattern theories from the practitioner literature, namely Auction Market Theory and Market Profile, without a peer-reviewed computational instantiation. We present ViperQ, a reinforcement learning system whose state representation is built explicitly from Auction Market Theory primitives: Volume Point of Control, Value Area position, Low Volume Node flags, Cumulative Volume Delta divergence, and tape-velocity signatures, assembled into a 20-dimensional Z-normalised vector. Two Proximal Policy Optimisation agents are trained with a prospect theory-grounded asymmetric reward function that penalises losing holds at a magnitude consistent with Kahneman and Tversky's loss-aversion coefficient. Evaluated on a held-out twelve-month partition of institutional tick data the agents have never seen, ViperQ achieves +163.6% ROI on TSLA (-27.5% max drawdown, 27,019 trades) and +116.5% ROI on NVDA (-47.8% max drawdown, 12,892 trades) under zero leverage. The results establish Auction Market Theory features as a tractable structured input modality for sequential decision-making on financial time series and motivate further work on microstructure-aware policy learning.

发表机构

  • University of Colorado, Colorado Springs(科罗拉多大学斯普林斯分校)
  • University of Iowa(爱荷华大学)

机构由 AI 辅助整理,请以论文原文为准。

↑