发表机构
California Institute of Technology; University of Bristol; Imperial College London; CERN(加州理工学院; 布里斯托大学; 伦敦帝国学院; 欧洲核子研究组织)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
Alkaid是一个开源领域专用编译器,通过算术中心IR和硬件感知优化,将亚微秒数据流内核快速转换为RTL/HLS代码,相比现有流程降低20-30% LUT使用和最多40%延迟,并支持非ML内核。
AI 中文摘要
在亚微秒级别运行的超低延迟机器学习和数据处理流水线通常包含静态数据流内核,这些内核需要细粒度的位宽控制、算术优化和快速的硬件性能估计。本文介绍了Alkaid,一个免费开源的领域专用编译器,可在数秒内将亚微秒延迟的数据流内核转换为平台无关的寄存器传输级(RTL)或高层次综合(HLS)就绪代码。Alkaid针对超低延迟机器学习流水线,尤其是那些由现有ML到硬件流程(如hls4ml)所处理的流水线,同时支持周围的预处理、后处理和控制逻辑。其核心是使用一种以算术为中心的中间表示——Alkaid低级IR(ALIR),以保留定点语义、异构位宽和操作级结构,用于优化和代码生成。Alkaid还提供了全面的硬件感知优化通道和一个可解释的白盒分析性能模型,用于无需综合循环的快速设计空间探索。在包括MLP、GNN和Transformer在内的代表性机器学习工作负载中,对于位精确等效设计,Alkaid的LUT使用量比先前最佳流程低20-30%,在某些情况下延迟降低高达40%。除了典型的神经网络,Alkaid在合成提升决策树方面达到了与最先进方法相当的性能,同时还支持实现和集成诸如排序网络和直方图等非机器学习内核。
英文摘要
Ultra-low-latency machine learning and data processing pipelines operating on the sub-microsecond level often contain static dataflow kernels that require fine-grained bitwidth control, arithmetic optimization, and fast hardware performance estimation. This works introduce Alkaid, a free and open source domain specific compiler that translates sub-microsecond latency dataflow kernels into platform agnostic Register Transfer Level (RTL) or High-Level Synthesis (HLS)-ready code in seconds. Alkaid targets ultra-low-latency ML pipelines, especially those addressed by existing ML-to-hardware flows such as \texttt{hls4ml}, while also supporting the surrounding pre-processing, post-processing, and control logic. At its core, Alkaid uses an arithmetic centric intermediate representation, Alkaid Low-level IR (ALIR), to preserve fixed-point semantics, heterogeneous bitwidths, and operation level structure for optimization and code generation. Alkaid further provides comprehensive hardware-aware optimization passes and an interpretable white-box analytical performance model for rapid design space exploration without synthesis in the loop. Across representative ML workloads, including MLPs, GNNs, and transformers, Alkaid achieves 20-30\% lower LUT usage than the best prior flows for bit-exact equivalent designs, with up to 40\% lower latency in some cases. Beyond typical neural networks, Alkaid achieves comparable performance to the state-of-the-art method for synthesizing boosted decision trees, while also enabling the implementation and integration of non-ML kernels such as sorting networks and histograms.