发表机构
Nara Institute of Science and Technology(奈良先端科学技术大学院大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究将BitNet b1.58推理映射到可编程CGLA上,通过可复用有符号int4乘加指令实现高效低比特LLM推理,FPGA测量频率缩放后每个乘积仅需0.390纳秒,实测吞吐达2.52 tokens/s。
AI 中文摘要
大语言模型(LLM)推理为每个生成的token传输模型权重和激活值,使得内存流量及其能量成本成为解码路径的一部分。BitNet b1.58用三元值表示其低比特权重,并使用整数激活值。然而,这种算术运算与传统的int8或浮点通用矩阵乘法不匹配,现有的BitNet加速器在专用数据通路中实现它。我们转而将此操作映射到CPU接地线性阵列(CGLA)上,这是一种可编程ASIC,具有显式直接内存访问、本地存储器和可复用的编译器可见整数通道。该映射添加了OP_SMA4作为可复用的有符号int4乘加指令,而非仅用于BitNet的数据通路。每个三元权重占用一个有符号4位通道。每个int8激活值被拆分为两个有符号int4片段,并通过移位相加进行重构。将145 MHz FPGA测量结果频率缩放至840 MHz 28 nm CGLA,每个有符号int4乘积耗时0.390纳秒。我们展示了CGLA卸载的BitNet C++执行测得2.52 tokens/s。
英文摘要
Large language model (LLM) inference transfers model weights and activations for every generated token, making memory traffic and its energy cost part of the decode path. BitNet b1.58 represents its low-bit weights by ternary values and uses integer activations. However, this arithmetic does not match conventional int8 or floating-point general matrix multiplication, and existing BitNet accelerators implement it in specialized datapaths. We instead map this operation to a CPU-Grounded Linear Array (CGLA), a programmable ASIC with explicit direct memory access, local memories, and reusable compiler-visible integer lanes. The mapping adds OP_SMA4 as a reusable signed-int4 multiply-accumulate instruction rather than a BitNet-only datapath. Each ternary weight occupies one signed 4-bit lane. Each int8 activation is split into two signed-int4 fragments and reconstructed by shift-and-add. Frequency scaling of the 145 MHz FPGA measurement to an 840 MHz 28 nm CGLA achieved 0.390 ns per signed-int4 product. We showed that CGLA-offloaded BitNet C++ execution measures 2.52 tokens/s.
CommentsAccepted as a Regular Paper for the CSA Workshop at CANDAR 2026