发表机构
Seoul National University; Hanyang University(首尔大学; 汉阳大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
Spexis通过投机并行与前瞻调度,在不增加KV缓存内存的前提下引入新并行维度,提升多GPU LLM推理效率,较最优基线加速达34%。
AI 中文摘要
Spexis是一个多GPU LLM推理框架,通过投机并行(speculative parallelism)提高流水线并行和张量并行的效率。Spexis并非仅将投机解码用于加速令牌生成,而是将投机与正常执行并行运行,引入新的并行维度,且不增加KV缓存内存使用。这提高了内存效率,并有助于缓解多GPU推理的瓶颈。Spexis进一步使用前瞻调度来预测投机质量和未来内存压力,从而减少浪费的投机、KV缓存驱逐和重新计算。基于vLLM构建,Spexis在多种GPU配置下大幅提升服务性能,与使用最优流水线和张量并行组合的基线相比,实现了高达34%的加速。Spexis的源代码可在此https URL公开获取。
英文摘要
Spexis is a multi-GPU LLM inference framework that improves the efficiency of pipeline and tensor parallelism through speculative parallelism. Rather than using speculative decoding only to accelerate token generation, Spexis runs speculation in parallel with normal execution, introducing a new parallelism axis without increasing KV-cache memory usage. This improves memory efficiency and helps mitigate the bottlenecks of multi-GPU inference. Spexis further uses lookahead scheduling to predict speculation quality and future memory pressure, allowing it to reduce wasted speculation, KV-cache eviction, and recomputation. Built on top of vLLM, Spexis largely improves serving performance across a range of GPU configurations, achieving speedups of up to 34% over a baseline that uses the optimal combination of pipeline and tensor parallelism. Spexis's source code is publicly available at https://github.com/mlsys-seo/spexis.
CommentsEMNLP 2026 main