发表机构
M2 Technologies(M2科技公司)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文通过四项扩展(fast_send、参与者组、可选邮箱队列、内存池)使参与者模型适配高频交易,实测表明同步往返仅数十纳秒,框架开销不足基础延迟的1%,并实现生产与回测代码统一。
AI 中文摘要
参与者模型——状态隔离、无数据竞争、抗死锁以及顺序单消息推理——长期以来被认为不适合高频交易(HFT):参与者似乎意味着许多线程、每个参与者一个邮箱、每次交互一个堆分配的消息以及一次上下文切换,这些开销与微秒级预算不相容。本文认为,对于同地部署的参与者,这种否定是错误的,并通过分析和已部署的实测实现(kaspar-hft,一个开源的C++20框架)来支持这一观点。四项扩展使该模型适应高频交易:fast_send,一种同步传递机制,发送线程内联运行接收者的处理程序并将回复作为值返回;参与者组,将参与者共同调度到一个线程上,共享一个邮箱;每个参与者可选择的邮箱队列;以及一个内存池。fast_send具有接收者透明性:处理程序无法判断传递是同步还是异步,也无法判断哪个线程运行它。分组同步链在一个线程上运行,将调度器上下文切换从O(N)减少到O(1),并且线程本地调用链测试在任何锁被获取之前捕获循环调用。微基准测试显示同步往返时间为数十纳秒。在实时CME市场数据流(ES、NQ、ZN期货)上,从套接字到订单簿的延迟分解为约7微秒的解码和订单簿基础延迟加上每条消息的斜率;框架自身的贡献不到基础延迟的1%。尾部不是由参与者机制决定的,而是由市场的非泊松、聚类到达过程决定的,这在配套论文中有所描述。共享队列组还产生了生产/模拟二元性:相同的参与者代码在实时交易和确定性回测中无需修改即可运行。
英文摘要
The actor model - state isolation, data-race freedom, and sequential single-message reasoning - has long been dismissed as unsuitable for high-frequency trading (HFT): actors seem to imply many threads, a mailbox per actor, and a heap-allocated message plus a context switch per interaction, overhead incompatible with a microsecond budget. This paper argues the dismissal is wrong for co-located actors, with a deployed, measured implementation: kaspar-hft, an open-source C++20 framework. Four extensions adapt the model for HFT: fast_send, a synchronous delivery mechanism in which the sending thread runs the receiver's handler inline and returns the reply as a value; actor groups, which co-schedule actors on one thread behind a shared mailbox; per-actor selectable mailbox queues; and a memory pool. fast_send has receiver transparency: the handler is written identically for synchronous and asynchronous delivery and does not depend on which was used or which thread runs it. A grouped synchronous chain runs on one thread, cutting scheduler context switches from O(N) to O(1); a thread-local call-chain test detects cyclic invocation on one thread before any lock is taken, while a cycle spread across threads deadlocks. Microbenchmarks put the synchronous round trip at tens of nanoseconds. On a live CME market-data feed (ES, NQ, ZN futures), when the socket-reader thread decodes each packet and updates the book itself, socket-to-book medians for book updates are 0.8-1.1 microseconds, and a fast_send hop is about 1% of that. The shared-queue group also yields a production/simulation duality: the same actor code runs unchanged in live trading and deterministic backtest.
Comments31 pages, 4 figures. Companion paper on the market-data arrival process in preparation