arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2610.03369q-fin.TRcs.LG

加密货币订单执行的专家混合模型:训练稳定性、尾部风险与失败模式

Mixture-of-Experts for Cryptocurrency Order Execution: Training Stability, Tail Risk, and Failure Modes

Alexander Ardaiz, Varun Budati, Ali Habibnia

首次发表
浏览论文内容

中文总结 AI 辅助

本研究在加密货币订单执行中评估专家混合模型,发现其未优于DDQL,且崩溃抑制源于训练规范而非架构,尾部风险随专家数增加而恶化。

中文摘要 AI 辅助

深度强化学习订单执行策略在不同训练种子下可能产生显著差异,因此表观上的架构改进可能反映的是有利的训练实现,而非架构本身可复现的特性。我们评估了标准双重深度Q学习(DDQL)、K均值分区的DDQL专家混合模型(K∈{2,4,8}),以及与K=4和K=8专家预算参数匹配的密集网络,使用来自Binance的5分钟均值聚合BTC/USDT限价订单簿数据。没有一种学习配置能显著改善相对于DDQL的平均实现缺口。在所报告的规范下,所有配置的平均缺口均高于TWAP(0.39个基点)和立即清算(0.21个基点),在该环境中,无摩擦回放和终端紧迫性惩罚使得早期清算几乎无成本;11/100次标准DDQL运行(而MoE K≥4组中均无)收敛到等待强制清算的策略。随后,我们在设备匹配的基线上分解该规范。仅退火探索就消除了观察到的崩溃(从12/100降至0/100;精确McNemar检验p=4.9×10⁻⁴),与专家分区下的消除效果相当。将退火探索与对齐奖励相结合,在19/30次运行中恢复了崩溃;加入全部三项规范变更后,崩溃率升至48/100。在该环境中,专家分区对于抑制崩溃并非必要,且似乎掩盖了训练规范失败而非赋予内在性能优势。在测试的六种规范下,MoE K=8运行均未崩溃。跨种子离散度在K=8时最低,但非单调且对族系调整不稳健,而策略内尾部风险随K单调恶化。失败模式的表观归因在30到100个种子之间发生反转,凸显了重复种子评估的重要性。

英文摘要

Deep reinforcement-learning policies for order execution can vary substantially across training seeds, so apparent architectural gains may reflect favourable training realisations rather than reproducible properties of the architecture. We evaluate vanilla Double Deep Q-Learning (DDQL), K-means-partitioned mixtures of DDQL experts at $K \in \{2, 4, 8\}$, and dense networks parameter-matched to the $K{=}4$ and $K{=}8$ expert budgets on 5-minute mean-aggregated BTC/USDT limit order book data from Binance. No learned configuration significantly improves mean implementation shortfall over DDQL. Under the reported specification, all have higher mean shortfall than TWAP (0.39 bps) and immediate liquidation (0.21 bps) in an environment whose frictionless replay and terminal-urgency penalty make early liquidation nearly costless; 11/100 vanilla-DDQL runs, versus none in either MoE $K{\geq}4$ arm, converge to a policy that waits until forced liquidation. We then decompose this specification on a device-matched baseline. Annealed exploration alone eliminates observed collapses (12/100 to 0/100; exact McNemar $p{=}4.9{\times}10^{-4}$), matching the elimination under expert partitioning. Combining annealed exploration with the aligned reward restores collapse in 19/30 runs; with all three specification changes, it rises to 48/100. In this environment, expert partitioning is unnecessary to suppress collapse and appears to mask a training-specification failure rather than confer an intrinsic performance benefit. No MoE $K{=}8$ run collapses under any of the six specifications tested. Across-seed dispersion is lowest at $K{=}8$ but non-monotone and not robust to family-wise adjustment, while within-policy tail risk worsens monotonically with $K$. The apparent attribution of the failure mode reverses between 30 and 100 seeds, illustrating the importance of repeated-seed evaluation.

发表机构

  • Virginia Tech(弗吉尼亚理工大学)
  • Dataism Laboratory for Quantitative Finance(数据主义量化金融实验室)

机构由 AI 辅助整理,请以论文原文为准。

↑