arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.26067cs.LG

FuncCode:在函数空间中压缩 Kolmogorov--Arnold 网络并采用硬件感知量化

FuncCode: Compressing Kolmogorov--Arnold Networks in Function Space with Hardware-Aware Quantization

Kazi Ahmed Asif Fuad, Lizhong Chen

首次发表
浏览论文内容

中文总结 AI 辅助

FuncCode 通过函数空间共享码本和硬件感知量化压缩 KAN,实现高压缩比并保持精度,显著降低 FPGA 内存。

中文摘要 AI 辅助

Kolmogorov--Arnold 网络(KANs)用可学习的单变量函数取代标量边权重,增加了灵活性,但也增加了参数内存,因为每条边存储多个系数,通常还带有单独的基础分支。我们引入了 FuncCode,一种与基无关的压缩方法,它从采样的边响应中形成共享码本,独立地对基和基础分支进行编码,并以量化的、位打包的格式导出生成的码本和每条边的索引。在样条和多项式 KANs 中,采样的边响应表现出比其系数表示低 13%--35% 的有效秩。进一步的重复对照表明,仅函数空间聚类在统计上与系数空间聚类相当;一致的精度提升来自于保留两个分支的不同共享结构。在十种子 MNIST 基准上,FuncCode 将样条和 GRAM KANs 压缩了 31.6 倍和 17.6 倍,精度损失仅为 0.31 和 0.34 个百分点。在具有 6.1M 边的卷积 KAGN 上,它实现了 19.9 倍的压缩,同时在 CIFAR-10 上保持与密集精度相差 0.54 个百分点以内,在 CIFAR-100 上相差 1.89 个百分点以内。压缩后,每条边的索引占存储权重位的比例高达 99.4%,使得表示受索引限制。在九个位精确的 FPGA 加速器上,FuncCode 将 SplineKAN 布线后权重内存相对于密集 INT4 减少了 3.87 倍,且不增加周期数或延迟。FuncCode 实现可在以下 https URL 获取。

英文摘要

Kolmogorov--Arnold Networks (KANs) replace scalar edge weights with learnable univariate functions, increasing flexibility but also parameter memory because each edge stores multiple coefficients, often together with a separate base branch. We introduce FuncCode, a basis-agnostic compression approach that forms shared codebooks from sampled edge responses, codes the basis and base branches independently, and exports the resulting codebooks and per-edge indices in a quantized, bit-packed format. Across spline and polynomial KANs, sampled edge responses exhibit $13$--$35\%$ lower effective rank than their coefficient representations. Further replicated controls show that function-space clustering alone is statistically tied with coefficient-space clustering; the consistent accuracy gain comes from preserving the distinct sharing structure of the two branches. On a ten-seed MNIST benchmark, FuncCode compresses spline and GRAM KANs by $31.6\times$ and $17.6\times$ with only $0.31$ and $0.34$ pp accuracy loss. On a 6.1M-edge convolutional KAGN, it achieves $19.9\times$ compression while remaining within $0.54$ pp of dense accuracy on CIFAR-10 and $1.89$ pp on CIFAR-100. After compression, per-edge indices account for up to $99.4\%$ of stored weight bits, making the representation index-bound. Across nine bit-exact FPGA accelerators, FuncCode reduces SplineKAN post-route weight memory by $3.87\times$ relative to dense INT4, without increasing cycle count or latency. The FuncCode implementation is available at https://github.com/OSU-STARLAB/FuncCode.

↑