AI 中文总结
研究针对大语言模型推理中令牌生成延迟瓶颈,提出无同步全规约算法SiFAR。通过消除一次性算法写后写依赖、利用交换机片上规约等方法,减少全规约延迟,提升端到端吞吐量,在特定条件下取得显著性能提升。
AI 中文摘要
推理模型和智能系统的兴起使大语言模型令牌生成延迟成为关键瓶颈。与聊天机器人不同,这些系统生成中间推理令牌,令牌延迟直接决定端到端响应时间。低延迟推理使用最小批处理,使令牌生成受带宽限制。张量并行通过跨GPU分片模型权重并并行加载来解决此问题,但扩展到更多GPU会引入随GPU数量增长的全规约开销。我们提出无同步全规约(SiFAR)算法,减少低延迟推理期间的同步开销。现有算法在通信前后存在屏障开销。我们通过协同设计通信和模型执行消除一次性算法中的写后写依赖,利用现代交换机的片上规约,提出冗余拉取和推测规约,减少全规约延迟。在TP=8时,SiFAR使Llama-3.1-8B的全规约延迟降低52%,端到端吞吐量提高18.6%,使Qwen3.5-397B-17B的端到端吞吐量提高13.1%。
英文摘要
The rise of reasoning models and agentic systems has made LLM token-generation latency a key bottleneck. Unlike chatbots, whose latency gains saturate at human reading speed, these systems generate intermediate reasoning tokens not consumed by humans. Thus, per-token latency directly determines end-to-end response time. Low-latency inference uses minimal batching, making token generation bandwidth-bound. Tensor Parallelism addresses this by sharding model weights across GPUs and loading them in parallel. However, scaling to more GPUs introduces All-Reduce overheads that grow with GPU count. Removing All-Reduce improves token throughput by 43% for Llama-3.1-8B on 8 H200 GPUs. We propose Synchronization-Free All-Reduce (SiFAR), which reduces synchronization overhead during low-latency inference. Existing oneshot and twoshot algorithms incur overheads from barriers before and after communication. First, we find that the bottom barrier in oneshot enforces a WAW dependency and eliminate it by co-designing communication and model execution to enable dual buffering. However, oneshot scales poorly with GPU count. Twoshot performs better at higher TP degrees but incurs an unavoidable bottom barrier. To overcome this, we leverage in-switch reduction in modern switches. We propose redundant pull, where each GPU reduces the full All-Reduce payload at the switch. This improves oneshot scalability while retaining its no-bottom-barrier advantage. Finally, to reduce top-barrier overhead, we observe that each decode step issues multiple All-Reduce operations, keeping GPUs tightly synchronized after the first. We therefore propose speculative reduction, which initiates data transfer before the top barrier and ensures correctness via lightweight validation. SiFAR reduces All-Reduce latency by up to 52% and improves end-to-end throughput by 18.6% for Llama-3.1-8B and 13.1% for Qwen3.5-397B-17B at TP=8.