arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.07740cs.DCcs.LGcs.PF

分布式Kolmogorov-Arnold网络训练在高性能计算系统上的可扩展性分析

Scalability Analysis of Distributed Kolmogorov-Arnold Network Training on High-Performance Computing Systems

Guangneng Chen, David Garcia Selfa, Pablo Quesada Barriuso

首次发表
浏览论文内容

中文总结 AI 辅助

本文实证研究了KAN在HPC系统上的数据并行训练可扩展性,发现其与传统深度学习一致,通信开销受All-Reduce影响,并提供部署指南。

中文摘要 AI 辅助

Kolmogorov-Arnold网络(KANs)用网络边上的可学习单变量函数替代多层感知器(MLPs)的固定激活函数和线性权重,提供了更好的可解释性,并在某些情况下具有竞争力的参数效率。虽然KANs的逼近性质已受到广泛关注,但它们在分布式、多GPU训练下的行为尚未被系统地表征。本文对多节点、多GPU高性能计算(HPC)基础设施上的数据并行KAN训练进行了实证可扩展性研究,从四个维度评估:强扩展、弱扩展、通信开销和模型规模扩展。实验在FinisTerrae III超级计算机上进行,使用最多8个NVIDIA A100 GPU,跨4个节点,采用PyTorch分布式数据并行(DDP)。KAN训练在8个GPU时达到74.7%的并行效率,加速比为5.97倍,与传统深度学习工作负载一致。弱扩展显示从单GPU到多GPU的初始吞吐量下降,随后保持强稳定性。通信开销呈现非单调模式(1.3%-6.1%),主要由All-Reduce算法选择和节点间延迟驱动,而非KAN的边级梯度结构。参数与内存之比随模型规模增大而改善,即使训练时间扩展不利。这些结果表明,针对KAN的算子级和数据并行优化是互补的。我们提供了GPU拓扑和模型规模选择的部署指南,并讨论了合成回归评估的局限性。

英文摘要

Kolmogorov-Arnold Networks (KANs) replace the fixed activation functions and linear weights of Multi-Layer Perceptrons (MLPs) with learnable univariate functions on network edges, offering improved interpretability and, in some settings, competitive parameter efficiency. While the approximation properties of KANs have received considerable attention, their behavior under distributed, multi-GPU training has not been systematically characterized. This paper presents an empirical scalability study of data-parallel KAN training on multi-node, multi-GPU high-performance computing (HPC) infrastructure, evaluated along four dimensions: strong scaling, weak scaling, communication overhead, and model-size scaling. Experiments were conducted on the FinisTerrae III supercomputer using up to 8 NVIDIA A100 GPUs across 4 nodes with PyTorch Distributed Data Parallel (DDP). KAN training reaches 74.7% parallel efficiency at 8 GPUs with a 5.97x speedup, consistent with conventional deep learning workloads. Weak scaling shows an initial single-to-multi-GPU throughput drop followed by strong stability. Communication overhead follows a non-monotonic pattern (1.3%-6.1%), driven primarily by All-Reduce algorithm selection and inter-node latency rather than KAN's edge-wise gradient structure. The parameter-to-memory ratio improves with model size even as training time scales unfavorably. These results indicate that operator-level and data-parallel optimizations for KAN are complementary. We provide deployment guidelines for GPU topology and model-size selection, and discuss the limitations of a synthetic-regression evaluation.

发表机构

  • University of Santiago de Compostela(圣地亚哥德孔波斯特拉大学)
  • Galicia Supercomputing Center (CESGA)(加利西亚超级计算中心(CESGA))

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑