arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.22275cs.DCq-bio.NC

非统一内存访问(NUMA)平衡阻碍脉冲神经网络模拟性能

NUMA balancing hampering performance of spiking network simulations

Melissa Lober, Alp Inangu, Gorka Peraza Coppola, Dennis Terhorst, Sebastian Gillessen, Jan Vogelsang, Hans Ekkehard Plesser, Brian Wylie, Benedikt Steinbusch, G… 展开作者

Melissa Lober, Alp Inangu, Gorka Peraza Coppola, Dennis Terhorst, Sebastian Gillessen, Jan Vogelsang, Hans Ekkehard Plesser, Brian Wylie, Benedikt Steinbusch, Guido Trensch, Susanne Kunkel, Markus Diesmann

AI总结:

研究发现关闭自动NUMA平衡可降能耗,其内存访问模式影响脉冲神经网络模拟性能分析。通过新性能显示揭示问题,发现自动NUMA平衡有弊端并影响相关库,还能检测系统扰动,为超级计算机提供用户级开关选项,助研究人员找最佳设置,该现象普遍性待研究。

AI中文摘要:

当今计算中心大多运行基于传统CPU和GPU的系统,降低能耗的直接方法是减少应用运行时间。神经形态计算有望提供提高人工智能能效的替代架构。在传统超级计算机上模拟大规模脉冲神经网络的代码是参考。关闭自动NUMA平衡可降低30%能耗,其内存访问模式与自动NUMA平衡动态交互,虽不影响模拟结果正确性,但在性能分析中影响时间测量波动。新的时间和计算节点解析性能显示揭示了分布式脉冲神经网络模拟的细粒度时间变异性,发现自动NUMA平衡有弊端并影响jemalloc库,该方法还能让开发者检测HPC系统扰动并针对性改进模拟技术。我们在超级计算机上为用户提供了按作业关闭自动NUMA平衡的选项,让研究人员找到适合应用的最佳设置,文献中有迹象表明该现象之前已被观察到,但在科学计算中似乎并非常识,其在科学代码中的普遍程度仍有待研究。

英文摘要:

Computing centers today mostly operate conventional CPU- and GPU-based systems, where the direct way of decreasing energy consumption is a reduction in the applications' runtime. Neuromorphic computing promises an alternative architecture with improved energy efficiency for artificial intelligence. In this endeavor, code for the simulation of large-scale spiking networks on conventional supercomputers is the reference. We show that turning off automatic NUMA balancing may reduce energy consumption by 30%. This dwarfs other attempts of increasing the energy efficiency of a computing center with respect to cost effectiveness. The memory access pattern of spiking network simulation code dynamically interacts with automatic NUMA balancing. This does not affect the correctness of simulation results and thus goes unnoticed in day-to-day neuroscience research. In performance analysis, however, time measurements fluctuate obstructing attempts to optimize simulation technology. A new time- and compute-node resolved performance display exposes the fine-grained temporal variability of distributed spiking network simulations. The analysis uncovers that automatic NUMA balancing is of disadvantage and affects the jemalloc library for thread-aware memory allocation in a transient manner. The method also allows developers to detect perturbations of the HPC system and target specific improvements to simulation technology. As a consequence, we have equipped our supercomputers with an option to turn on or off automatic NUMA balancing on a per-job basis on the user level. This gives researchers the opportunity to find the best setting for the application at hand. There are indications in the literature that the effect has been observed before, yet it does not seem common knowledge in scientific computing. It remains to be investigated how widespread the phenomenon is among scientific codes.

↑