arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

Kolmogorov-Arnold网络的可学习激活函数的神经缩放定律与演化

Neural scaling laws and evolution of learnable activation functions of Kolmogorov-Arnold networks

Tilen Cadez, Sanghoon Lee, Kyoung-Min Kim

arXiv 2610.00985首次发表:更新:

发表机构

Asia Pacific Center for Theoretical Physics; Pohang University of Science and Technology(亚太理论物理中心; 浦项科技大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究探讨KANs的神经缩放定律及激活函数随数据量增加的演化,发现断裂缩放行为,并为高效应用KANs提供定量指导。

AI 中文摘要

Kolmogorov-Arnold网络(KANs)代表了传统基于多层感知器(MLP)的神经网络的一种引人注目的替代方案。通过将激活函数作为可学习元素,KANs提供了优越的可解释性,使其适用于科学领域。在这项工作中,我们研究了KANs的神经缩放定律及其可学习激活函数在数据集扩展下的结构演化。具体来说,我们评估了三种KAN变体——BSRBF-KAN、Gottlieb-KAN和Faster-KAN——在标准图像分类基准(MNIST和Fashion-MNIST)以及一个专门的科学回归任务(从莫尔磁性纹理的域图像中估计磁参数)上的缩放行为。我们的结果表明,测试损失${\cal L}$表现出一种断裂的神经缩放定律(BNSL)行为,作为数据集大小$N_D$的函数。在通过随机猜测区域后,损失遵循架构和任务相关的缩放行为。对于图像分类任务,损失从较快缩放分支过渡到较慢缩放分支,${\cal L}\propto N_D^{-\alpha}$和${\cal L}\propto N_D^{-\beta}$,其中$\alpha>\beta$。指数$\alpha$和$\beta$强烈依赖于特定的网络架构和数据集大小区域,分别从0.4到1.5和从0.06到0.6。对于磁参数回归任务,损失遵循单一缩放定律,其指数从1.28到2.59。此外,我们提供了激活函数如何随着数据量的增加而细化其复杂性的结构分析,发现数据集扩展驱动了从简单的线性近似向稳定、可解释的符号形式的转变。这些发现为有效应用KANs同时管理模型表达性与计算开销之间的权衡提供了定量路线图。

英文摘要

Kolmogorov-Arnold Networks (KANs) represent a compelling alternative to traditional Multi-Layer Perceptron (MLP)-based neural networks. By employing activation functions as learnable elements, KANs offer superior interpretability, making them suited for scientific domains. In this work, we investigate the neural scaling laws of KANs and the structural evolution of their learnable activation functions under dataset expansion. Specifically, we evaluate the scaling behavior of three KAN variants---BSRBF-KAN, Gottlieb-KAN, and Faster-KAN---across standard image classification benchmarks (MNIST and Fashion-MNIST) and a specialized scientific regression task (magnetic parameter estimation from domain images of moiré magnetic textures). Our results demonstrate that the test loss ${\cal L}$ exhibits a broken neural scaling law (BNSL) behavior as a function of the dataset size $N_D$. After passing through a random-guess regime, the loss follows architecture- and task-dependent scaling behavior. The loss crosses from a faster- to a slower-scaling branch, ${\cal L}\propto N_D^{-α}$ and ${\cal L}\propto N_D^{-β}$ with $α>β$ for image classification tasks. The exponents $α$ and $β$ depend strongly on both the specific network architecture and the dataset-size regime, ranging from 0.4 to 1.5 and from 0.06 to 0.6, respectively. For the magnetic parameter-regression task, the loss follows a single scaling law with its exponent ranging from 1.28 to 2.59. Additionally, we provide a structural analysis of how activation functions refine their complexity as data volume increases, finding that dataset expansion drives a transition from simple linear-like approximations toward stable, interpretable symbolic forms. These findings provide a quantitative roadmap for the efficient application of KANs while managing the trade-off between model expressivity and computational overhead.

Comments13 pages, 10 figures. Supplementary Notes will be provided in the published version

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑