查找表无损压缩用于硬件应用
Lossless Compression of Lookup Tables for Hardware Applications
浏览论文内容
中文总结 AI 辅助
本文提出CompressedLUT,一种用于硬件查找表的无损压缩方案,结合分解、自相似性、高位压缩和多级压缩技术,在FPGA上实现高效存储,适用于非线性函数、常系数乘法器和KANs。
中文摘要 AI 辅助
大型查找表在硬件中被广泛用于存储常量值数组,应用范围从基本数学运算(如常系数乘法和非线性函数评估)到新兴的机器学习模型(包括基于表格的神经网络(NNs)和Kolmogorov-Arnold网络(KANs))。然而,存储大量常量值表格可能导致资源受限的边缘设备(如FPGA)中过高的硬件成本。在本文中,我们提出了CompressedLUT,一种无损压缩方案及其解码器硬件架构,用于在硬件中高效存储和检索任意数据。我们的方法结合了分解、自相似性、高位压缩和多级压缩技术,以最大化表格大小节省且不损失精度。其硬件解码器主要使用加法、算术右移和几个小型查找表,确保低面积和高吞吐量。我们在FPGA上通过实现多个非线性函数、常系数乘法器(CCMs)和KANs(12位分辨率)评估了CompressedLUT。CompressedLUT作为开源工具提供。
英文摘要
Large lookup tables are widely used in hardware to store constant-valued arrays for applications ranging from elementary mathematical operations, such as constant-coefficient multiplication and nonlinear function evaluation, to emerging machine learning models, including table-based neural networks (NNs) and Kolmogorov-Arnold networks (KANs). However, storing extensive tables of constant values can lead to excessive hardware costs in resource-constrained edge devices such as FPGAs. In this paper, we propose CompressedLUT, a lossless compression scheme and its decoder hardware architecture for the efficient storage and retrieval of arbitrary data in hardware. Our method combines decomposition, self-similarities, higher-bit compression, and multilevel compression techniques to maximize table size savings without accuracy loss. Its hardware decoder primarily uses addition, arithmetic right shift, and several small lookup tables, ensuring low area and high throughput. We evaluated CompressedLUT on FPGAs by implementing multiple nonlinear functions, constant-coefficient multipliers (CCMs), and KANs at 12-bit resolution. CompressedLUT is available as an open-source tool.
发表机构
- University of Minnesota(明尼苏达大学)
机构由 AI 辅助整理,请以论文原文为准。