TEMPO:面向内存密集型与计算密集型场景的感知完成时间的专家并行负载均衡
TEMPO: Makespan-Aware Expert-Parallel Load Balancing Across Memory- and Compute-Bound Regimes
浏览论文内容
中文总结 AI 辅助
TEMPO是一种感知完成时间的专家并行调度器,可在内存与计算密集型混合场景下高效平衡负载,在8-GPU测试中优于现有方案,提升了MoE模型的吞吐量并降低了延迟。
中文摘要 AI 辅助
在专家混合(MoE)模型的专家并行(EP)部署中,每一层的同步速度由最慢的GPU决定。现有的调度器平衡令牌数量(如EPLB、LPLB、UltraEP)或激活专家数量(如METRO),均假设专家处理时间与其中一个变量呈线性关系。对两代数据中心GPU的测量显示,专家处理时间与这两个变量均非简单线性关系:当令牌数低于$\nstar\approx156$--$168$时,高带宽内存(HBM)权重流传输占主导,成本与激活的副本数量相关,而非令牌数;当令牌数高于该阈值时,分组通用矩阵乘法(GEMM)将令牌打包为128令牌的M块,因此拆分专家会增加填充后的计算量。我们提出了一个最大仿射模型$t=\max(a+bG,\\,c+\beta N)$来描述这两种场景。实际解码批次中,热门专家处于线性场景,冷门专家同时处于平坦场景;实测批次显示,代理调度的建模块时间差异为1.4--1.6倍(第95百分位最高达1.7倍),且哪种代理更优会随场景变化。我们将每批次调度形式化为固定费用的完成时间问题——在两个完全复制的GPU上为NP难问题,在退化极限下为多项式问题——并提出了TEMPO,一种感知完成时间的调度器,可在关键路径外以毫秒级解决该问题;其与SGLang的集成在进程外运行,将调度与计数收集融合为一个图内内核。基于8个GPU的Testbed A微基准测试,TEMPO在所有场景下均与最佳固定基准的差距在1%以内,在场景混合时最高提升15.5%。在Testbed B上的端到端测试中,Qwen3-235B(处于优势区域)的吞吐量提升4--6%,第99百分位延迟降低约15.6%;DeepSeek-V3(处于外部、通信主导区域)仅显示机制成本。我们的结论是,不存在通用的最优方案,而应采用相图来预测部署前的两种结果。
英文摘要
In expert-parallel (EP) MoE serving, every layer synchronizes at the slowest GPU. Dispatchers balance token counts (EPLB, LPLB, UltraEP) or activated-expert counts (METRO), assuming expert time is linear in one. Measurements on two datacenter GPU generations show it is neither: below $n^* \approx 156$--$168$ tokens, HBM weight streaming dominates---cost attaches to $activated replicas$, not tokens; above it, grouped GEMM rounds tokens to 128-tile $M$-tiles, so $splitting$ an expert adds padded compute. A max-affine profile $t=\max(a+bG,\,c+βN)$ captures both regimes. Realistic decode batches hold hot experts in the linear regime and cold in the flat $simultaneously$; recorded batches show proxy dispatches differ by $1.4$--$1.6\times$ in modeled block time (p95 up to $1.7\times$), and $which$ proxy wins flips with the regime. We formalize per-batch dispatch as a fixed-charge makespan problem---NP-hard on two fully replicated GPUs, polynomial in degenerate limits---and present TEMPO, a makespan-aware dispatcher solving it in milliseconds off the critical path; its SGLang integration runs out-of-process and fuses dispatch with count collection into one in-graph kernel. Anchored by an 8-GPU Testbed A microbenchmark, TEMPO stays within $1\%$ of the best fixed baseline everywhere and wins by up to $15.5\%$ where regimes mix. End-to-end on Testbed B, Qwen3-235B (inside the win region) gains $4$--$6\%$ throughput and cuts p99 latency by $\sim 15.6\%$; DeepSeek-V3 (outside, communication-dominated) shows only mechanism cost. A phase diagram, not a universal win, is the claim: it predicts both outcomes before deployment.
发表机构
- KlingAI Research(KlingAI研究院)
机构由 AI 辅助整理,请以论文原文为准。