发表机构
LG Electronics; University of Toronto(LG电子; 多伦多大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对卸载权重导致的LLM解码慢问题,提出SpAx方法,通过将权重读取分为保留、近似和省略三档,利用激活稀疏性,在CPU和闪存卸载下分别实现最高5.57倍和4.81倍加速,且困惑度增加不超过10%。
AI 中文摘要
在显存不足以容纳权重的消费级GPU上部署LLM,可能导致推理速度慢得难以接受,因为解码过程中需要反复将卸载的权重从系统RAM或闪存以远低于本地GPU显存访问的带宽传输到GPU。激活稀疏性通过跳过与零或接近零激活相关的权重来减少这些传输。然而,随着省略的激活贡献增多,模型质量最终会迅速下降,这表明与小幅度激活相关的权重共同对模型质量有显著影响。在本工作中,我们改善了利用激活稀疏性时模型质量与解码性能之间的权衡。我们的核心思想是将是否读取权重的二元选择替换为三种选项:完全保留、使用压缩权重表示近似、或完全省略。SpAx跳过与最接近零的激活相关的权重,读取较小幅度激活的近似权重,并读取最大幅度激活的原始权重。较小幅度的激活会减弱近似权重引入的误差,而压缩权重表示需要传输的字节数更少。当权重卸载到CPU内存时,SpAx在16位权重下平均加速解码3.86倍(最高5.57倍),在4位权重下平均加速2.06倍(最高2.74倍),同时WikiText-2困惑度增加最多10%。当权重卸载到闪存时,加速分别为平均3.31倍(最高4.81倍)和1.54倍(最高2.03倍)。
英文摘要
Deploying LLMs on consumer-grade GPUs with insufficient memory to hold their weights can result in prohibitively slow inference, because decoding repeatedly transfers offloaded weights from system RAM or flash storage into GPU at much lower bandwidth than local GPU-memory access. Activation sparsity reduces these transfers by skipping weights associated with zero or near-zero activations. However, as more activation contributions are omitted, model quality eventually degrades rapidly, indicating that weights associated with small-magnitude activations collectively influence model quality sharply. In this work, we improve the trade-off between model quality and decoding performance when exploiting activation sparsity. Our key idea is to replace the binary choice of whether or not to read a weight with three options: fully retain it, approximate it using a compressed weight representation, or omit it entirely. SpAx skips weights associated with activations closest to zero, reads approximate weights for smaller-magnitude activations, and reads original weights for the largest-magnitude activations. Smaller-magnitude activations attenuate the errors introduced by approximate weights, while compressed weight representations require fewer bytes to be transferred. With weights offloaded to CPU memory, SpAx speeds up decoding by 3.86X on average (up to 5.57X) with 16-bit weights and 2.06X (up to 2.74X) with 4-bit weights, at a WikiText-2 perplexity increase of at most 10%. With weights offloaded to flash storage, the speedups are 3.31X on average (up to 4.81X) and 1.54X (up to 2.03X).