发表机构
Shahid Beheshti University(沙希德·贝赫什提大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
RecKAN提出基于可学习二阶多项式递推的新型KAN基,在多分类、时间序列预测等任务上优于参数匹配的KAN基线及MLP,且基的递推系数具有可解释性。
AI 中文摘要
柯尔莫哥洛夫-阿诺尔德网络(KAN)将标准网络的固定标量权重替换为每条边上的可学习单变量函数,但现有变体仍固定了这些函数所基于的“基”,如B样条、切比雪夫多项式、小波或雅可比多项式,仅学习基上的组合权重。我们提出RecKAN,其通过二阶多项式递推定义基本身:$R_{n+1}(x) = (ax^2+bx+c)R_n(x) + (dx+e)R_{n-1}(x)$,该递推式的5个系数与网络共同学习。我们证明此递推式可恢复多个经典多项式族,包括两类切比雪夫多项式、斐波那契多项式、佩尔多项式和雅各布斯塔尔多项式作为特例,并证明在包含所有这些多项式的子族上,其次数随n线性增长,这为可学习基能突破任何固定经典选择提供了具体依据。在涵盖图像、文本、生物医学时间序列分类及时间序列预测的多个基准数据集上,RecKAN在所有分类任务中均优于3个参数匹配的KAN基线(切比雪夫、雅可比和样条基),并在ETTh1预测基准上取得最低均方误差(MSE)。此外,当与卷积骨干网络结合作为分类头时,RecKAN在Fashion MNIST、CIFAR-10和SVHN上的准确率高于标准MLP头。在合成函数拟合基准上,它能拟合参数相当的MLP欠拟合的剧烈振荡目标函数。我们进一步表明,学习到的递推系数具有可解释性:在需要最多局部结构的任务上,训练会使基偏离包含我们识别出的所有经典族的线性次数增长机制,这与我们对该结构转变作用的理论分析一致。
英文摘要
Kolmogorov--Arnold Networks (KANs) replace the fixed scalar weights of a standard network with learnable univariate functions on each edge, but existing variants still fix the \emph{basis} that those functions are built from: B-splines, Chebyshev polynomials, wavelets, or Jacobi polynomials, and learn only the combination weights over it. We introduce RecKAN, which instead defines the basis itself by a second order polynomial recurrence, $R_{n+1}(x) = (ax^2+bx+c)R_n(x) + (dx+e)R_{n-1}(x)$, whose five coefficients are learned jointly with the network. We show this recurrence recovers several classical polynomial families including both kinds of Chebyshev polynomials, Fibonacci, Pell, and Jacobsthal polynomials as special cases, and prove that its degree grows linearly in $n$ exactly on the sub-family containing all of them, giving a concrete sense in which the learned basis can move beyond any fixed classical choice. Across multiple benchmark datasets spanning image, text, biomedical time series classification, and time series forecasting, RecKAN outperforms three parameter-matched KAN baselines (Chebyshev, Jacobi, and spline based) on all classification tasks and achieves the lowest MSE on the ETTh1 forecasting benchmark. Additionally, when used as a classifier head with a convolutional backbone, RecKAN achieves higher accuracy than standard MLP heads on Fashion MNIST, CIFAR-10, and SVHN. On a synthetic function fitting benchmark it tracks a sharply oscillatory target that a parameter comparable MLP under fits. We further show that the learned recurrence coefficients are interpretable: on the task requiring the most local structure, training moves the basis away from the linear degree growth regime that contains every classical family we identify, consistent with our theoretical analysis of what that structural shift enables.
Comments21 pages , 7 figures