arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

几何约束柯尔莫哥洛夫-阿诺德网络:通过巴拿赫对偶学习边几何

Geometry-Constrained Kolmogorov-Arnold Networks: Learning Edge Geometry via Banach Duality

K S Sesh Kumar

arXiv 2608.25807首次发表:更新:

发表机构

Brevan Howard Centre for Financial Analysis; Imperial Business School(布赖恩·霍华德金融分析中心; 帝国理工商学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究提出几何约束KANs,通过每条边的标量指数p学习边几何,在50个符号回归任务中,其在噪声下稳定性优于多数固定基模型,小样本表现更优,可学习指数具可解释性。

AI 中文摘要

柯尔莫哥洛夫-阿诺德网络(KANs)将深度架构中的固定激活函数替换为可学习的单变量边函数,因此边参数化的选择至关重要。现有变体依赖固定基函数,如样条、多项式或傅里叶特征,这些在观测数据前就已确定函数空间几何。我们提出几何约束KANs,这是一类从巴拿赫对偶映射导出的边激活函数,其中几何本身通过每条边的标量指数p>1进行学习。该指数控制定性响应:亚欧几里得值产生类似ℓ₁(LASSO)几何的尖锐阈值类行为;p=2时恢复线性机制;更大值在原点附近产生更平缓的响应。在50个符号回归目标(40个来自AI Feynman基准,加上10个合成压力测试)上,几何约束KANs在中位数归一化均方根误差(NRMSE)上匹配或击败所有固定基基线(Banach-KAN为0.030,与切比雪夫基线持平且优于样条基线);在平均排名上,Banach-KAN在18个方程核心任务中排名最佳(2.00),在完整基准中与最强样条基线统计持平(2.32对2.34)。在测量噪声下增益最明显:当σ从0增至1时,ℓᵖ-KAN仅下降3.7倍,甚至低于交叉验证样条(约11倍),而未正则化样条下降21.6倍;Banach-KAN下降8.8倍,与调优后的样条相当,但比未正则化样条稳定得多。Banach-KAN在小样本 regime中也获得最多的单方程胜利,仅当训练集增大时固定基模型才会追上。学习到的指数提供可解释的相对信号:在固定初始化下,它们揭示了跨方程族和输入维度的一致、依赖目标的几何排序。

英文摘要

Kolmogorov-Arnold Networks (KANs) replace fixed activations in deep architectures with learnable univariate edge functions, making the choice of edge parametrisation central. Existing variants rely on fixed bases such as splines, polynomials, or Fourier features, which impose a function-space geometry before data are observed. We introduce geometry-constrained KANs, a family of edge activations derived from Banach duality maps in which the geometry itself is learned through a scalar exponent $p > 1$ per edge. This exponent controls the qualitative response: sub-Euclidean values produce sharp, threshold-like behaviour reminiscent of the $\ell_1$ (LASSO) geometry, $p = 2$ recovers the linear regime, and larger values produce flatter responses near the origin. Across 50 symbolic-regression targets ($40$ from the AI Feynman benchmark plus $10$ synthetic stress tests), geometry-constrained KANs match or beat every fixed-basis baseline on median NRMSE (Banach-KAN $0.030$, tying Chebyshev and improving on splines); on average rank Banach-KAN is best on the $18$-equation core ($2.00$) and statistically tied with the strongest spline on the full benchmark ($2.32$ vs. $2.34$). The clearest gains appear under measurement noise: as $σ$ grows from $0$ to $1$, $\ell^p$-KAN degrades only $3.7\times$ -- below even a cross-validated spline ($\approx 11\times$) -- while an unregularised spline degrades $21.6\times$; Banach-KAN degrades $8.8\times$, comparable to a tuned spline but far more stable than the unregularised one. Banach-KAN also takes the most per-equation wins in the small-sample regime, with fixed-basis models catching up only as the training set grows. Learned exponents provide an interpretable, relative signal: at a fixed initialisation they reveal a consistent, target-dependent geometric ordering across equation families and input dimensions.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑