面向BrainScaleS神经形态系统芯片化扩展的统一互连网络
A Unified Interconnection Network for Chiplet-Based Scaling of the BrainScaleS Neuromorphic System
- Institute of Computer Engineering, Heidelberg University(海德堡大学计算机工程研究所)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本文提出一种用于BrainScaleS-2神经形态系统芯片化扩展的路由芯片设计,通过统一互连网络在2D网格中高效处理容错脉冲与不容错数据,仿真验证了高带宽利用率。
AI中文摘要:
BrainScaleS-2(BSS-2)神经形态架构将脉冲神经网络(SNN)原语的模拟仿真与紧密耦合的ADC和数字处理单元相结合。这些模拟SNN原语是固定的硬件资源,无法进行复用,从而将仿真网络规模限制为物理硬件副本的数量。为了克服模拟设计扩展的挑战,基于芯片的设计提供了一种有前景的方法,在成本和灵活性方面优于单片扩展。实现基于芯片的BSS-2架构需要一种互连网络来处理两类不同的流量:容错流量(如脉冲)和不容错流量(如配置数据或处理单元交换的数据)。本工作提出了BSS-2架构的路由芯片设计,使得多个BSS-2单元能够在2D网格拓扑中互连。两类流量通过单个宽并行芯片间链路进行复用。利用SNN的容错性,脉冲以非安全和同步的方式传输,接收侧的到达时间直接决定脉冲的突触前时间。相反,不容错的数据传输通过点对点自动重传请求协议进行保护,并使用基于信用的流控。这些设计选择在仿真中得到验证,所提出的设计在$1 \ imes 10^{-10}$的误码率下,对于替代梯度训练用例,在广泛的脉冲与安全流量比率范围内,可以维持$95\%$的链路带宽利用率,且对脉冲时间抖动影响很小。
英文摘要:
The BrainScaleS-2 (BSS-2) neuromorphic architecture combines analog emulation of spiking neural network (SNN) primitives with tightly coupled ADCs and digital processing units. These analog SNN primitives are fixed hardware resources that cannot be multiplexed, limiting the emulated network size to the number of physical hardware copies. To overcome the challenges of scaling analog designs, chiplet-based designs offer a promising approach with cost and flexibility advantages over monolithic scaling. Implementing a chiplet-based BSS-2 architecture requires an interconnection network that handles two distinct classes of traffic: error-tolerant traffic such as spikes and error-intolerant traffic like configuration data or data exchanged by the processing units. This work presents the design of a routing chiplet for the BSS-2 architecture that enables interconnection of multiple BSS-2 units in a 2D mesh topology. Both traffic classes are multiplexed over a single wide parallel die-to-die link. Exploiting the fault tolerance of SNNs, spikes are transmitted unsecured and synchronously, with the arrival time on the receiving side directly determining the pre-synaptic time of the spike. Conversely, error-intolerant data transmission is secured by a point-to-point Automatic Repeat Request protocol and uses credit-based flow control. These design choices are validated in simulation, where the proposed design can sustain $95\,\%$ link bandwidth utilization for the use case of surrogate gradient training across a wide range of spike-to-secured traffic ratios under a $1 \times 10^{-10}$ bit error rate with little impact on spike timing jitter.