arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

数据包、交易与队列:基于CME市场数据测量研究的HFT系统设计原则

Packets, Transactions and Queues: Design Principles for HFT Systems from a Measurement Study of CME Market Data

Vincent Maciejewski

arXiv 2609.32848首次发表:更新:

发表机构

M2 Tech(M2科技)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

通过测量CME市场数据,验证了HFT系统单线程事件循环规则,提出基于数据包到达与交易时间关系的队列设计原则,并给出分析框架支持。

AI 中文摘要

HFT系统传统上被构建为单线程事件循环,其依据是每次线程跳转都会增加延迟。我们通过一项对NQ近月合约超过一年的CME市场数据的测量研究来检验这一规则,追踪了数据流中两个交易所时间戳所标记的每个数据包和匹配引擎交易,并将结果与一个实时生产接收器进行了核对。数据包以近乎临界的自激聚类形式到达,这些聚类属于匹配引擎的交易,而非交易所打包数据包的方式。引擎通常在一微秒的几分之一内处理连续交易,而市场数据发布器每个发布周期(约7.5微秒)最多发送一个数据包,因此突发流量会以间隔一个周期的数据包序列到达接收器。这为HFT系统提供了设计原则。第一,如果接收器在每个发布周期内处理一个数据包,则无论市场如何突发,都不会在到达时排队;此时单线程是最佳选择。第二,超过该周期后,会出现由交易时间而非数据包速率或大小驱动的排队尾部,此时两个线程可能优于一个线程:将服务链拆分为两个阶段并置于不同线程上,可以消除大部分尾部延迟,但代价是中位数延迟增加一跳。第三,只有最慢的阶段才重要,因此拆分只有在缩短该阶段时才有价值。第四,在略低于该周期(生产接收器运行之处)的情况下,剩余的尾部延迟来自多消息数据包和可变服务时间,此时的关键杠杆是每条消息的成本和服务时间的分布,而非线程数量。一个分析框架、一个突发限制吞吐量恒等式以及将串联系统精确简化为单个瓶颈服务器,为这些结果提供了支持。

英文摘要

HFT systems are conventionally built as a single-threaded event loop, on the rule that every thread hop adds latency. We test that rule against a measurement study of more than a year of CME market data for the NQ front-month contract, following every packet and matching-engine transaction through the feed's two exchange timestamps, and checking the results against a live production receiver. Packets arrive in near-critical self-exciting clusters that belong to the matching engine's transactions, not to how the exchange packs them. The engine often processes consecutive transactions within a fraction of a microsecond, while the market-data publisher sends at most one packet per publisher period of about 7.5 microseconds, so a burst reaches the receiver as a train of packets one period apart. This yields design principles for HFT systems. First, a receiver that handles each packet within one publisher period never queues on arrivals, however bursty the market; there one thread is best. Second, above that period a queueing tail appears, driven by the timing of transactions, not by packet rate or size, and two threads can be better than one: splitting the servicing chain into two stages on separate threads removes most of the tail at the cost of one hop on the median. Third, only the slowest stage matters, so a split pays only if it shortens it. Fourth, just under the period, where the production receiver runs, the remaining tail comes from multi-message packets and variable service times, and the levers are cost per message and spread of service, not thread count. An analytic framework, a burst-limit throughput identity and an exact reduction of the tandem to a single bottleneck server, supports these results.

Comments123 pages, 36 figures, 31 tables. Code in the Kaspar-HFT repository (https://github.com/vincent212/kaspar-hft)

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑