arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

SW-KAN:基于Stieltjes-Wigert q-正交多项式的Kolmogorov-Arnold网络

SW-KAN: Kolmogorov-Arnold Networks with Stieltjes-Wigert q-Orthogonal Polynomials

Amirhosein Azarpour, Seyyed Moein Kazemi

arXiv 2610.00050首次发表:更新:

发表机构

Shahid Beheshti University(沙希德·贝赫什提大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出SW-KAN,采用Stieltjes-Wigert q-正交多项式解决KAN中域不匹配问题,通过平滑映射和高效递推实现优越的精度-效率权衡,在资源受限条件下表现稳健。

AI 中文摘要

Kolmogorov-Arnold网络(KAN)通过将固定的节点激活替换为边上的可学习单变量函数,代表了深度学习中的范式转变,提供了更强的可解释性和参数效率。尽管近期基于多项式的KAN变体解决了原始B样条实现的计算开销问题,但它们引入了一个基本但尚未充分探索的挑战:无界实值输入与正交多项式基的有界或半无限支撑之间的域不匹配。为解决这一限制,我们提出了Stieltjes-Wigert Kolmogorov-Arnold网络(SW-KAN),这是一种新颖的架构,采用定义在半无限域(0, infinity)上的Stieltjes-Wigert q-正交多项式。我们引入了一种平滑的指数-双曲正切映射,在保持良好条件梯度的同时稳定地弥合域差距,并利用数值稳定的三项递推关系,在O(N)次操作中无需特殊函数调用即可评估多项式展开。通过涵盖图像分类和连续函数逼近的综合实验,我们证明了SW-KAN在多样化任务中实现了优越的精度-效率权衡。Stieltjes-Wigert多项式的对数正态权重结构和可学习的q参数提供了一种独特的归纳偏置,使得在资源受限条件下(包括降低的特征维度和有限的训练数据)也能实现稳健的性能。所提出的架构不仅在标准基准上优于已建立的多项式KAN基线,而且在用极少数参数逼近复杂多元函数时表现出强大的表示能力,使其成为资源受限环境中高效函数逼近和分类的有力替代方案。

英文摘要

Kolmogorov-Arnold Networks (KANs) represent a paradigmatic shift in deep learning by replacing fixed node activations with learnable univariate functions on edges, offering enhanced interpretability and parameter efficiency. While recent polynomial-based KAN variants have addressed the computational overhead of original B-spline implementations, they introduce a fundamental yet underexplored challenge: the domain mismatch between unbounded real-valued inputs and the bounded or semi-infinite support of orthogonal polynomial bases. To address this limitation, we propose the Stieltjes-Wigert Kolmogorov-Arnold Network (SW-KAN), a novel architecture that employs Stieltjes-Wigert q-orthogonal polynomials defined on the semi-infinite domain (0, infinity). We introduce a smooth exponential-of-tanh mapping that stably bridges the domain gap while preserving well-conditioned gradients, and leverage a numerically stable three-term recurrence that evaluates polynomial expansions in O(N) operations without special-function calls. Through comprehensive experiments spanning image classification and continuous function approximation, we demonstrate that SW-KAN achieves superior accuracy-efficiency trade-offs across diverse tasks. The log-normal weight structure and learnable q-parameter of Stieltjes-Wigert polynomials provide a distinct inductive bias that enables robust performance under resource-constrained conditions, including reduced feature dimensionality and limited training data. The proposed architecture not only outperforms established polynomial KAN baselines on standard benchmarks but also exhibits strong representational capacity for approximating complex multivariate functions with remarkably few parameters, making it a compelling alternative for efficient function approximation and classification in resource-constrained settings.

Comments22 pages, Code and pretrained models available at: https://github.com/amirhoseinazarpour/SW-KAN

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑