arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.23841cs.ARcs.LGcs.PF

流水线原生Transformer:为带宽高效的自回归解码协同设计模型架构与CPU推理

Pipeline-Native Transformers: Co-Designing Model Architecture and CPU Inference for Bandwidth-Efficient Autoregressive Decode

Tom Poperszky

首次发表
浏览论文内容

中文总结 AI 辅助

该研究协同设计流水线原生Transformer架构与CPU推理引擎cflow,在TinyStories训练任务上实现了更低的权重带宽、更少的缓存缺失,且解码速度优于同类模型的CPU后端。

中文摘要 AI 辅助

CPU上的单token自回归解码受内存带宽而非算术运算限制:现代CPU的计算能力约为1 TFLOP/s,但主内存带宽仅约50 GB/s,且每个生成的token必须一次性读取所有活跃权重。本报告认为最有效的应对方式是协同设计模型架构与推理运行时,提出了一款CPU优先的流引擎cflow,以及一系列流水线原生Transformer架构,其层间依赖图被构建为允许采用垂直的、以阶段为主的执行调度。cflow以计算消耗顺序将权重存储为L2大小的块,仅读取每个混合专家(MoE)层的前k个专家,融合投影操作,并根据每个模型的依赖参数执行感知延迟的调度。在基于TinyStories训练的五种架构中,其中一种(arch2_4_combined)实现了关键路径权重带宽降低2.00倍(从9.00 MB/token降至4.50 MB/token),且其困惑度仅与最佳候选模型相差0.24;该块布局相比行主序基线,L1数据读取缺失减少7.29倍。在32 vCPU的Ice Lake服务器上,309亿参数的流水线原生MoE模型通过cflow实现解码速度为5.94 token/s,优于同等规模密集模型的this http URL(4.75 token/s)和vLLM CPU后端(1.65 token/s)。将专家延迟窗口实现为磁盘驻留专家层上的异步I/O重叠,可进一步获得最高1.68倍的净增益,与重叠模型的性能差距在1%以内。测量结果推翻了八项设计主张中的一项,且第二项结论不明确,两项均完整报告了其成立的条件。

英文摘要

Single-token autoregressive decode on CPUs is bound by memory bandwidth, not arithmetic: a modern CPU sustains roughly 1 TFLOP/s of compute but only about 50 GB/s from main memory, and each generated token must stream every active weight once. This report argues that the most effective response is to co-design the model architecture and the inference runtime together. It presents cflow, a CPU-first streaming engine, alongside a family of pipeline-native transformer architectures whose inter-layer dependency graphs are constructed to permit a vertical, stage-major execution schedule. cflow stores weights as L2-sized tiles in compute-consumption order, reads only the top-k experts of each mixture-of-experts layer, fuses projections, and executes a delay-aware schedule from per-model dependency parameters. Across five architectures trained on TinyStories, one (arch2_4_combined) achieves a 2.00x reduction in critical-path weight bandwidth (9.00 to 4.50 MB/token) within 0.24 perplexity of the best candidate, and the tile layout incurs 7.29x fewer L1-data read misses than a row-major baseline. On a 30.9-billion-parameter pipeline-native MoE, cflow decodes at 5.94 tokens/s (tok/s) on a 32-vCPU Ice Lake server, ahead of llama.cpp (4.75) and the vLLM CPU backend (1.65) on comparably sized dense models. Realizing the expert-delay window as asynchronous I/O overlap on a disk-resident expert tier yields a further net win of up to 1.68x, matching the overlap model within 1%. Measurement refutes one of the eight design claims and leaves a second inconclusive; both are reported in full, with the conditions under which they would hold.

补充信息

↑