发表机构
POSTECH; ISTA(浦项科技大学; 奥地利科学技术研究所)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对MoE部署中的内存和带宽瓶颈,提出软硬件协同框架,通过连续重参数化实现联合稀疏量化,并定制分组稀疏GEMM内核,在万亿参数规模下提升精度并显著加速推理。
AI 中文摘要
混合专家(Mixture-of-Experts, MoE)架构使得前沿语言模型能够扩展到万亿参数规模,但其部署受到巨大内存占用和内存带宽限制的制约。尽管现代加速器提供了稀疏张量核(Sparse Tensor Cores, SpTCs),通过低精度半结构化稀疏性来减少权重存储并提高吞吐量,但由于显著的模型质量下降和缺乏分组稀疏GEMM原语,在MoE中利用这些特性仍然具有挑战性。我们提出了一种端到端的软硬件协同设计框架,该框架将专家权重压缩为硬件原生的低精度稀疏表示,并在SpTCs上加速其执行。在算法层面,我们的框架通过连续重参数化放宽了离散的半结构化支撑选择,使得在路由器加权重建目标下,能够与量化权重进行可微分的联合优化,并实现可扩展的专家并行压缩。在系统层面,我们开发了一个定制的分组稀疏GEMM内核,专门针对SpTCs上的低精度稀疏MoE推理进行优化。在从300亿到一万亿参数的MoE模型上,我们的框架将最先进的联合稀疏量化精度提高了最多4.35个百分点,同时保留了原始模型96.09%的性能。在NVIDIA B200 GPU上,我们的内核性能比供应商基线高出最多1.65倍,将服务吞吐量提高了1.18倍,并将端到端延迟降低了最多4.03倍。这些结果确立了软硬件协同设计作为实现可扩展且高效MoE部署的实用路径。
英文摘要
Mixture-of-Experts (MoE) architectures allow frontier language models to scale to trillions of parameters, but their deployment is constrained by massive memory footprints and memory-bandwidth limitations. Although modern accelerators provide Sparse Tensor Cores (SpTCs) that reduce weight storage and increase throughput through low-precision semi-structured sparsity, exploiting them for MoEs remains challenging because of substantial model-quality degradation and the lack of grouped sparse GEMM primitives. We present an end-to-end hardware-software co-design framework that compresses expert weights into hardware-native, low-precision sparse representations and accelerates their execution on SpTCs. Algorithmically, our framework relaxes discrete semi-structured support selection through continuous reparameterization, enabling differentiable joint optimization with quantized weights under a router-weighted reconstruction objective and scalable expert-parallel compression. Systemically, we develop a custom grouped sparse GEMM kernel tailored to low-precision sparse MoE inference on SpTCs. Across MoE models ranging from 30 billion to one trillion parameters, our framework improves state-of-the-art joint sparse-quantization accuracy by up to 4.35 percentage points while preserving 96.09% of the original model's performance. On NVIDIA B200 GPUs, our kernel outperforms the vendor baseline by up to $1.65\times$, increasing serving throughput by $1.18\times$ and reducing end-to-end latency by up to $4.03\times$. These results establish hardware-software co-design as a practical path toward scalable and efficient MoE deployment.