AI 中文总结
针对GPU张量计算流水线的固定并行度与粗粒度调度瓶颈,提出扩展SIMT模型的FIBER架构,通过线程-寄存器解耦实现动态并行缩放与细粒度调度,在LLM服务场景下获显著加速。
AI 中文摘要
现代GPU越来越多地将Tensor Cores集成到执行流水线中。尽管得益于从Ampere架构的基于寄存器的操作数供应到Hopper和Blackwell架构的无冗余、基于内存的操作数供应,总张量吞吐量持续增长,但为现代AI工作负载高效编排完整的张量计算流水线仍然具有挑战性。我们确定根本瓶颈为固定并行度和粗粒度调度,这两者在现代AI工作负载中暴露出来,此类工作负载将各种非GEMM操作与GEMM操作交织在一起。为了高效编排张量计算,我们提出FIBER,这是一种扩展GPU SIMT(单指令多线程)模型的新架构。其基本执行实例fiber与私有寄存器所有权解耦,仅携带最小控制状态,同时通过共享视图访问SM的寄存器。这实现了动态并行度缩放、细粒度寄存器级数据流调度,并为矩阵操作数供应提供了无冗余的替代方案。我们扩展ISA、微架构和编译器以实现共享寄存器寻址、无冲突操作数交付以及基于fiber的程序映射。在典型的混合精度LLM服务场景下,FIBER在Ampere上实现了2.25倍的端到端加速(原始FP16计算的加速为1.15倍),在Hopper和Blackwell上分别实现了1.8倍和2.09倍的加速,内核级增益高达2.49倍。
英文摘要
Modern GPUs increasingly integrate Tensor Cores into the execution pipeline. Although aggregate tensor throughput continues to grow, aided by an operand supply that has evolved from register-based in Ampere to redundancy-free, memory-based in Hopper and Blackwell, efficiently orchestrating the complete tensor compute pipeline for the modern AI workloads remains challenging. We identify the fundamental bottlenecks as fixed parallelism and coarse-grained scheduling, both of which are exposed by modern AI workloads that interleave diverse non-GEMM operations with GEMM. To orchestrate tensor computation efficiently, we propose FIBER, a new architecture that extends the GPU SIMT (single instruction, multiple thread) model. Its basic execution instance, the \emph{fiber}, is decoupled from private register ownership, carrying only minimal control state while accessing an SM's registers through a shared view. This enables dynamic parallelism scaling, fine-grained register-level dataflow scheduling, and offers a redundancy-free alternative for matrix operand supply. We extend the ISA, microarchitecture, and compiler to realize shared-register addressing, conflict-free operand delivery, and fiber-based program mapping. Under a typical mixed-precision LLM serving scenario, FIBER achieves a 2.25x end-to-end speedup on Ampere (1.15x for the original FP16 computation), with 1.8x and 2.09x on Hopper and Blackwell respectively, and kernel-level gains up to 2.49x.