arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

面向英特尔AI PC集群的分布式大语言模型推理的预编译流水线分片

Pre-Compiled Pipeline Shards for Distributed LLM Inference on Intel AI PC Fleets

Tate Berenbaum, Muthaiah Venkatachalam

arXiv 2608.19147首次发表:更新:

发表机构

Not Community Labs Inc.; Intel Corporation(非社区实验室公司; 英特尔公司)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出基于预编译流水线分片的分布式LLM推理方案,借助OpenVINO优化、投机解码与微批处理,实现AI PC集群高效运行单设备无法承载的大模型,提升了吞吐量与交互速度。

AI 中文摘要

现代英特尔AI PC配备了集成GPU和NPU,拥有16GB以上的统一内存,且存在大量空闲时间,但该内存不足以容纳700亿参数大语言模型这类大型模型。本文展示了少量AI PC通过普通网络协同工作,可运行单台设备无法承载的模型。研究采用流水线并行技术:按层将模型拆分为各阶段分片,每个分片预编译为OpenVINO图,每台设备运行一个分片并将激活值传递给下一分片。三项技术使该方案足够实用:一是恢复未分片模型的速度,原生按阶段导出的模型因缺失OpenVINO GPU优化(间接键值缓存融合)而远低于整体推理速度,在每个分片中注入beam_idx gather操作可触发该优化,使分片性能达到与整体模型相当的水平;二是利用有状态OpenVINO模型的投机解码;三是通过在各阶段交错处理多个用户的请求(每个请求携带自身缓存,即微批处理),让流水线同时为多个用户服务。实验结果显示:双节点Llama 3.1 8B INT4流水线在相同硬件上为两个并发用户服务时,吞吐量是未分片模型单用户吞吐量的1.79倍,在模拟广域延迟下差距会进一步扩大;该设计可扩展至单台设备无法承载的70B模型,英特尔Tiber云上的四台Lunar Lake AI PC部署可为单用户提供交互级速度,且输出token与无投机的四节点流水线解码完全一致。代码、原始基准测试日志及复现脚本作为独立包存放在该https URL的顶层reproduction/目录中。

英文摘要

Modern Intel AI PCs ship capable integrated GPUs and NPUs with 16+ GB of unified memory, and they spend considerable time idle. That is not enough memory to fit a large model such as a 70B-parameter LLM. We show that a handful of AIPCs, working together over an ordinary network, can serve models beyond the capability of any single one. We use pipeline parallelism: a model is split by layer into per-stage shards, each pre-compiled into an OpenVINO graph, so that every machine runs one shard and passes activations to the next. Three techniques make this fast enough to be useful. First, we recover the speed of the unsplit model: a naive per-stage export runs well below monolithic inference because it misses an OpenVINO GPU optimization, and injecting a beam_idx Gather into each shard triggers that optimization (the IndirectKVCache fusion) and brings the shards to parity. Second, we leverage speculative decoding on stateful OpenVINO models. Third, the pipeline serves several users at once by interleaving their requests across the stages, each request carrying its own cache (micro-batching). Together, a two-node Llama 3.1 8B INT4 pipeline serves two concurrent users at 1.79x the single-user throughput of the unsplit model on the same hardware, and the gap widens under simulated wide-area latency. The same design scales to a 70B model that no single fleet member can hold: a four-node deployment of Lunar Lake AI PCs on Intel Tiber Cloud serves a single user at interactive speed, with output token-for-token identical to the same four-node pipeline decoding without speculation. Code, raw benchmark logs, and reproduction scripts ship as a self-contained package at https://github.com/labscommunity/pipeline-sharded-inference-paper (in the top-level reproduction/ directory).

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑