SparseKAN:在基函数、神经元和比特维度上压缩柯尔莫哥洛夫-阿诺德网络(KAN)
SparseKAN: Compressing Kolmogorov--Arnold Networks Across Basis Functions, Neurons, and Bits
浏览论文内容
中文总结 AI 辅助
SparseKAN是一种在基函数、神经元、数值精度维度压缩KAN的方法,通过分层可学习门硬化重要性结构,可大幅减少参数、降低延迟并保持准确率,提升软硬件效率。
中文摘要 AI 辅助
柯尔莫哥洛夫-阿诺德网络(KAN)用由多个基函数系数参数化的可学习单变量函数替代了标量边权重,这引入了传统神经网络压缩无法直接发现的冗余来源。我们提出SparseKAN,这是一种在三个互补维度上压缩KAN的统一方法:基函数、神经元/通道和数值精度。SparseKAN为基础分支、非线性基分支和单个基项配备了分层可学习门,这些门在可微分的激活-代价目标下进行训练。随后,在明确的基函数和宽度预算下,将学习到的重要性结构硬化,以全精度或低精度恢复,并物理压缩为更小的稠密张量,而非保留为稀疏掩码。在MNIST、CIFAR-10和CIFAR-100上,针对样条、多项式、RBF、小波和卷积KAN变体的实验表明,这些结构维度在代价上可预测地组合。我们还发现项重要性存在强烈的基函数依赖差异:在评估的Gram-多项式设置中,基于系数的选择比匹配的低阶截断最多高出15.25个准确率点。8位量化具有广泛的鲁棒性,而4位卷积KAN需要量化感知适配。物理压缩在MNIST上可移除多达73.0%的参数且无准确率损失,并将大批次CUDA延迟降至稠密执行的0.51倍。在ZCU104 FPGA上,得到的稀疏低位模型实现了高达23.63倍的推理延迟降低,证明SparseKAN将功能冗余转化为可衡量的软件和硬件效率。SparseKAN的实现可在该https URL获取。
英文摘要
Kolmogorov--Arnold Networks (KANs) replace scalar edge weights with learnable univariate functions parameterized by multiple basis coefficients. This introduces a source of redundancy that conventional neural-network compression does not directly expose. We present \textbf{SparseKAN}, a unified approach that compresses KANs along three complementary axes: basis functions, neurons/channels, and numerical precision. SparseKAN equips the base branch, nonlinear basis branch, and individual basis terms with hierarchical learnable gates trained under a differentiable active-cost objective. The learned importance structure is subsequently hardened under explicit basis and width budgets, recovered in full or low precision, and physically compacted into smaller dense tensors rather than retained as sparse masks. Experiments on MNIST, CIFAR-10, and CIFAR-100 across spline, polynomial, RBF, wavelet, and convolutional KAN variants show that the structural axes compose predictably in cost. We also find strong basis-dependent differences in term importance: coefficient-based selection outperforms matched low-order truncation by up to 15.25 accuracy points in the evaluated Gram-polynomial settings. Eight-bit quantization is broadly robust, whereas 4-bit convolutional KANs require quantization-aware adaptation. Physical compaction removes up to 73.0\% of parameters without accuracy loss on MNIST and reduces large-batch CUDA latency to as little as $0.51\times$ dense execution. On a ZCU104 FPGA, the resulting sparse low-bit models achieve up to $23.63\times$ lower inference latency, demonstrating that SparseKAN converts functional redundancy into measurable software and hardware efficiency. The SparseKAN implementation is available at https://github.com/OSU-STARLAB/SparseKAN.