arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

用于小型语言模型的柯尔莫哥洛夫-阿诺德网络

Kolmogorov--Arnold Networks for Small Language Models

Felippe Alves, Renato Vicente

arXiv 2607.15525首次发表:更新:

AI 中文总结

研究探讨柯尔莫哥洛夫-阿诺德网络(KANs)能否替代变压器前馈网络。通过在特定KAN中重建边缘、修剪等测试其特性,还评估多种网络在BabyLM等上的表现。结果表明小基KANs可用于审计标量变换,但替代方案未展现优于MLP基线的一致优势。

AI 中文摘要

柯尔莫哥洛夫-阿诺德网络(KANs)用学习到的一维边缘函数取代固定节点激活,为解释提供了明确接口,可能替代变压器前馈网络。我们分别测试这些说法。在一个六层、1000万参数的B样条KAN中,我们重建了所有884736个前馈边缘:87.8%超过(NLS>0.1)且0.4%不活跃。修剪最低活跃度的20%-25%导致损失增加可忽略不计,尽管结构化MLP神经元修剪能容忍相当的稀疏度。审计在BabyLM上重复,但网格大小扫描表明接近完全的fPCA压缩和高封闭形式拟合覆盖率是低容量网格2基的属性,而非通用KAN行为。对于替代,我们在BabyLM上评估了MLP、SwiGLU、分组切比雪夫和有理GR-KAN网络。KAN家族和门控变体在验证损失上优于GELU MLP,但这种排序在标准化基准测试中不适用:在十个种子和59875个BLiMP对中,准确率在62.4%-63.1%之间,EWoK仍处于随机水平,GR-KAN对BLiMP的(+0.7)分效应在补充测试中反转。更大规模测试也有警示:参数匹配的MLPEdge在Wikitext-103上表现不如MLP,286万参数的GR-KAN在稳定后仍低于SwiGLU ClimbMix基线。因此,小基KANs为审计学习到的标量变换提供了实用的、可语料库转移的接口,但测试的替代方案在与强大的MLP基线相比时,没有显示出一致的基准、质量或延迟优势。

英文摘要

Kolmogorov--Arnold Networks (KANs) replace fixed node activations with learned one-dimensional edge functions, offering an explicit interface for interpretation and a possible alternative to transformer feed-forward networks. We test these claims separately. In a six-layer, 10M-parameter B-spline KAN, we reconstruct all 884,736 feed-forward edges: 87.8\% exceed (NLS>0.1) and 0.4\% are inactive. Pruning the lowest-activity 20--25\% causes negligible loss increase, although structured MLP neuron pruning tolerates comparable sparsity. The audit replicates on BabyLM, but grid-size sweeps show that near-total fPCA compression and high closed-form-fit coverage are properties of the low-capacity grid-2 basis, not universal KAN behavior. For replacement, we evaluate MLP, SwiGLU, grouped Chebyshev, and rational GR-KAN networks on BabyLM. The KAN-family and gated variants improve validation loss over the GELU MLP, but this ordering does not transfer to standardized benchmarks: across ten seeds and 59,875 BLiMP pairs, accuracies span 62.4--63.1\%, EWoK remains at chance, and a (+0.7)-point GR-KAN effect on BLiMP reverses on the supplement. Larger tests are also cautionary: parameter-matched MLPEdge underperforms the MLP on Wikitext-103, and 286M-parameter GR-KAN remains below a SwiGLU ClimbMix baseline after stabilization. Thus, small-basis KANs provide a practical, corpus-transferable interface for auditing learned scalar transformations, but the tested replacements show no consistent benchmark, quality, or latency advantage over strong MLP baselines.

Comments30 pages, 11 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑