发表机构
Intel Corporation(英特尔公司)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对三元大语言模型存储效率问题,提出基于零权重分布的自适应布局BITCOS,将存储成本降至每权重1.485比特,并在多种CPU和GPU上实现最高1.28倍加速。
AI 中文摘要
三元大语言模型(LLM)将每个权重存储为三个符号$\{-1,0,+1\}$之一,因此三元模型的成本通常参照信息论下限$\log_2 3 \approx 1.585$比特/权重。当前主流的部署格式将五个三元权重打包进一个字节(五进制三值打包),由于实际使用的组大小为2的幂次,这导致每个权重向上取整为$1.625$比特。这种有效存储位宽将三个符号$\{-1,0,+1\}$视为等概率。我们测量了29个三元LLM模型的实际符号分布,发现零权重占比高达$51.5\\%$。受此发现启发,我们提出了BITCOS,一种简单的分布自适应布局,由密集的存在性位图和压缩的符号向量组成,在模型权重零密度为$z$时,每个权重元素成本为$2 - z$比特。在29个测试模型中,BITCOS在26个模型上比五进制三值打包存储更紧凑,并在最稀疏的模型上达到每权重$1.485$比特。BITCOS适用于现代处理器和GPU上的高效解包,我们为AVX-512、AVX2和Intel Xe2 GPU提供了优化的解包序列。与生产级最先进的三元矩阵-向量乘法内核相比,在真实世界三元模型展现的零密度下,我们提出的布局实现的收益最高达$1.28\times$。最后,我们展示了在5个不同平台(客户端和服务端CPU、集成和独立Xe2 GPU)上的端到端LLM推理结果,其中CPU上的解码吞吐量提升最高达$1.18\times$,GPU上最高达$1.27\times$。
英文摘要
Ternary Large Language Models (LLM) store every weight as one of three symbols $\{-1,0,+1\}$, so the cost of a ternary model is conventionally referenced to the information-theoretic $\log_2 3 \approx 1.585$ bits per weight. The prevailing deployment format packs five ternary weights into one byte (five-trit packing), and due to the power-of-two group sizes used in practice this rounds up to $1.625$ bits per weight. This effective storage bit-width treats the three symbols $\{-1,0,+1\}$ as equiprobable. We measure the actual symbol distribution of 29 ternary LLM models and find that zeros account for up to $51.5\%$ of all weights. Motivated by this finding, we introduce BITCOS, a simple distribution-adaptive layout comprised of a dense presence bitmap plus a compacted sign vector, and costs $2 - z$ bits per weight element given a zero density $z$ in the model's weights. BITCOS stores weights more compactly than the five-trit packing in 26 of the 29 tested models, and reaches $1.485$ bits per weight on the sparsest of them. BITCOS is amenable to efficient unpacking on modern processors and GPUs, and we present optimized unpacking sequences for AVX-512, AVX2 and Intel Xe2 GPUs. Measured against production state-of-the-art ternary matrix-vector multiplication kernels, at the zero densities real-world ternary models exhibit, the realized gain with our proposed layout is up to $1.28\times$. Finally, we illustrate end-to-end LLM inference results on 5 different platforms (client and server CPUs, integrated and discrete Xe2 GPUs) where decode throughput improves by up to $1.18\times$ on CPUs and $1.27\times$ on GPUs.