模型铸造与低参数门控:迈向更稀疏激活的前馈网络
Model Casting and Low-Parameter Gating: Towards More Sparsely Activated FFNs
浏览论文内容
中文总结 AI 辅助
本文提出模型铸造与LoPA门控,通过稀疏化FFN激活减少计算,实现高达3.2倍FLOP加速及GPU上3.31倍实际加速,显著优于现有方法。
中文摘要 AI 辅助
本文介绍了模型铸造(model casting),一种能够大幅稀疏化前馈网络(FFN)层内激活的中间训练(mid-training)方法。采用此策略后,在推理时,我们首先计算门控矩阵的输出,并利用其高稀疏性,避免与另外两个矩阵的计算,从而将FLOP计数减少高达3倍。虽然这一理论加速是上限,但模型铸造在CPU和GPU上均能转化为显著的加速效果。接着,我们引入了LoPA门控(LoPA Gating),一种新的FFN设计,可提高最大理论加速比。它是对门控矩阵的一种低FLOP参数化,通过相比另外两个稀疏激活的FFN矩阵,分配更少的FLOP和参数给门控矩阵,从而突破3倍上限。我们考虑了两种情况:(i)使用具有稀疏诱导激活的预训练模型进行铸造;(ii)从头开始使用LoPA进行训练。在所有设置中,我们都显著优于现有的剪枝解决方案和常规的RELU化(RELU-fication)。例如,在匹配质量下,使用LoPA铸造(LoPA Casting)我们实现了3.2倍的FLOP加速,而竞争方法top-p和TEAL最佳仅为1.6倍。使用专用内核,我们在90%稀疏度下实现了GPU上实际的3.31倍加速,超过了标准门控的3倍上限;与此同时,RELU化在低于80%稀疏度时趋于平稳。
英文摘要
This paper introduces model casting, a mid-training recipe that drastically sparsifies the activations within the Feed-Forward Network (FFN) layer. With this strategy, at inference time, we first compute the output of the gating matrix and, thanks to its high sparsity, we avoid computations with the two other matrices, reducing the FLOP count by up to 3x. While this theoretical speedup is an upper bound, model casting translates into significant speedups both on CPU and GPU. We then introduce LoPA Gating, a new FFN design that increases the maximum theoretical speedup. It is a low-FLOPs parameterization of the gating matrix that overcomes the 3x cap by allocating fewer FLOPs and parameters to the gating matrix, compared to the two other FFN matrices that are sparsely activated. We consider two cases: (i) we cast a pre-trained model with a sparsity inducing activation; (ii) we train with LoPA from scratch. In all settings, we significantly outperform existing pruning solutions and regular RELU-fication. For instance, at matched quality, we achieve a 3.2x FLOP speedup with LoPA Casting, against 1.6x at best for competing methods top-p and TEAL. Using dedicated kernels, we achieve an actual 3.31x speed-up on GPU at 90% sparsity, past the 3x ceiling of standard gating; RELU-fication, meanwhile, plateaus below 80% sparsity.