发表机构
The University of Hong Kong; Harbin Institute of Technology(香港大学; 哈尔滨工业大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究针对二值量化LLM推理的硬件内核缺失问题,提出FluxBin算法-内核协同设计,实现超低位推理的加速、能效提升与内存压缩,可高效部署70B规模模型。
AI 中文摘要
尽管二值量化在理论上可为大语言模型(LLM)提供极致压缩与加速,但现有研究常忽视专用硬件内核的必要性,因持续依赖昂贵的浮点运算或运行时反量化开销,未能释放全部加速潜力。为填补该空白,我们提出FluxBin(Flexible LUT-based Ultra-low-bit eXecution with Binary bases,即基于灵活查找表的超低位二值基执行方案),这是一种算法-内核协同设计,将后训练量化与高度优化的CUDA内核相结合。算法层面,我们引入解耦行列二值分解以提升表征能力,同时保持硬件效率,辅以Hessian引导的显著性感知混合基以保留关键信息;内核层面,我们实现了带尺度融合的查找表构建方法以减少浮点运算,采用虚拟列映射将不规则、稀疏且具显著性的矩阵转换为适合密集执行的形式。大量评估表明,FluxBin在各类模型架构上实现了最高5.92倍的加速与10.19倍的能效提升,精度与经大量微调的方法相当,可在单块A100 GPU上部署700亿参数规模的模型,内存减少4倍。代码可访问该https URL获取。
英文摘要
While binary quantization theoretically promises extreme compression and acceleration for Large Language Models (LLMs), existing research often overlooks the necessity of specialized hardware kernels, thus failing to unleash the full acceleration potential due to persistent reliance on expensive floating-point arithmetic or runtime dequantization overheads. To bridge this gap, we propose FluxBin (\textbf{F}lexible \textbf{L}UT-based \textbf{U}ltra-low-bit e\textbf{X}ecution with \textbf{Bin}ary bases), an algorithm-kernel co-design that synergizes post-training quantization with a highly optimized CUDA kernel. Algorithmically, we introduce Decoupled Row-Column Binary Decomposition to enhance representational capacity while maintaining hardware efficiency, complemented by a Hessian-guided saliency-aware hybrid bases that preserve critical information. At the kernel level, we implement a Lookup Table Building Approach with Scale Fusion to reduce floating-point arithmetic, featuring a Virtual Columnar Mapping that transforms irregular, sparse, and salient matrices into dense execution. Extensive evaluations demonstrate FluxBin achieves up to $5.92\times$ speedup and $10.19\times$ energy savings across diverse model architectures, delivering comparable accuracy to heavily fine-tuned methods. This effectively enables the deployment of 70B-scale models on one single A100 GPU with a $4\times$ memory reduction. Code is available at https://github.com/nicyyyy/FluxBin.