arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2610.08403cs.LG

SSR:用于三元GEMM加速的稀疏段归约

SSR: Sparse Segment Reduction for Ternary GEMM Acceleration

Adeline Pittet, Shien Zhu, Valérie Verdan, Gustavo Alonso

首次发表
浏览论文内容

中文总结 AI 辅助

本文提出SSR方法,通过专用三元数据格式和利用稀疏模式的计算树,加速三元LLM推理,在45-95%稀疏度下比RSR++快2.1-11.3倍,并实现端到端3.5-6.3倍加速。

中文摘要 AI 辅助

大型语言模型(LLMs)需要大量的计算资源,限制了它们在资源受限硬件上的部署。三元LLMs通过三元值进行权重量化来缓解这些需求,实现了显著的压缩,通常具有50-90%的稀疏度。然而,现有方法存在局限性:针对三元权重优化的方法,如BitNet、冗余段归约(RSR)及其改进版本RSR++,并未利用稀疏结构,而传统稀疏格式忽略了三元特性,从而失去了双重优化机会。在本文中,我们提出了稀疏段归约(SSR),一种三元矩阵乘法方法,旨在加速三元LLMs和一般三元权重网络(TWNs)的推理。SSR具有专门优化的三元数据格式和一种算法,通过随稀疏度扩展的计算树系统地利用稀疏模式。SSR在稀疏度超过50%时提供了理论上的增益,其推理速度渐近快于RSR++,而实际评估显示在所有稀疏度水平上性能均有提升。评估结果表明,在45-95%稀疏度的三元GEMM上,SSR相比RSR++实现了2.1-11.3倍的加速。此外,在Llama-3 1B模型推理上,SSR相比RSR++实现了3.5-6.3倍的端到端加速和4.9%的内存节省。

英文摘要

Large Language Models (LLMs) require substantial computational resources, limiting their deployment on resource-constrained hardware. Ternary LLMs mitigate these demands through weight quantization via ternary values, achieving significant compression often with 50-90% sparsity. However, existing approaches have limitations: methods optimized for ternary weights, such as BitNet, redundant segment reduction (RSR), and its improved version RSR++, do not exploit sparsity structures, while conventional sparse formats neglect ternary characteristics, foregoing dual optimization opportunities. In this paper, we introduce Sparse Segment Reduction (SSR), a ternary matrix multiplication method designed to accelerate the inference of ternary LLMs and general Ternary Weight Networks (TWNs). SSR has a dedicated optimized ternary data format and an algorithm that systematically exploits sparsity patterns through computation trees that scale with the sparsity. SSR provides theoretical gains with asymptotically faster inference than RSR++ for sparsity above 50%, while practical evaluations reveal performance improvements across all sparsity levels. Evaluation results show that SSR achieves 2.1-11.3x speedup over RSR++ on ternary GEMM with 45-95% sparsity. Furthermore, SSR achieves 3.5-6.3x end-to-end speedup and 4.9% of memory saving over RSR++ on the Llama-3 1B model inference.

发表机构

  • ETH Zurich(苏黎世联邦理工学院)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑