arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

用于FPGA推理的带流水线读出的时间复用脉冲神经网络加速器

A Time-Multiplexed Spiking Neural Network Accelerator with Pipelined Readout for FPGA Inference

Reza Ansari, Maciej Wielgosz

arXiv 2608.00595首次发表:更新:

AI 中文总结

本文设计实现了带流水线读出的SNN加速器,在FPGA上针对MNIST分类优化,提升了工作频率与能效,为实时神经形态边缘推理提供了资源高效方案。

AI 中文摘要

脉冲神经网络(SNN)通过离散时间事件处理信息,是传统人工神经网络的高能效类神经形态替代方案。本文提出并在现场可编程门阵列(FPGA)上实现了一种仅用于推理的SNN加速器,针对MNIST数字分类任务进行了优化。为解决低成本器件固有的物理布线约束与时序瓶颈,我们设计了优化的硬件微架构,包括由有限状态机(FSM)控制的时间复用1比特脉冲馈送机制、用于权重存储的局部分布式内存,以及选择寄存器宽度以防止溢出的基于整数的泄漏整合发放(LIF)神经元模型。此外,多周期流水线化的argmax(最大值)及平局决胜读出模块消除了主要的组合关键路径。该加速器基于入门级AMD Artix-7 FPGA(XC7A200T)实现,采用784-64-10的网络拓扑,所提出的流水线架构将最大工作频率(Fmax)从13.3 MHz提升至167 MHz。硬件评估显示,每张图像的顺序处理延迟为82 μs,使1000个样本的VHDL仿真批处理可在0.082 s内完成。Vivado实现后基于矢量的功耗分析估计,总片上功耗为0.336 W,能效约为36300样本/焦耳。这些结果表明,所提出的微架构为实时神经形态边缘推理提供了资源高效的解决方案,前提是网络规模保持在时间复用执行的实际限制范围内。

英文摘要

Spiking Neural Networks (SNNs) provide a power-efficient neuromorphic alternative to traditional artificial neural networks by processing information through discrete temporal events. This paper presents the design and Field-Programmable Gate Array (FPGA) implementation of an inference-only SNN accelerator optimized for MNIST digit classification. To address the physical routing constraints and timing bottlenecks inherent in low-cost devices, we propose an optimized hardware microarchitecture featuring a time-multiplexed 1-bit spike-feeding mechanism governed by a finite state machine (FSM), localized distributed memory for weight storage, and an integer-based Leaky Integrate-and-Fire (LIF) neuron model with register widths selected to prevent overflow. In addition, a multi-cycle pipelined argmax and tie-breaker readout module eliminates the dominant combinational critical path. Implemented on an entry-level AMD Artix-7 FPGA (XC7A200T) using a 784-64-10 network topology, the proposed pipelined architecture increases the maximum operating frequency (Fmax) from 13.3 MHz to 167 MHz. Hardware evaluation demonstrates a sequential processing latency of 82 μs per image, enabling a 1,000-sample VHDL simulation batch to be completed in 0.082 s. Vivado post-implementation vector-based power analysis estimates the total on-chip power consumption at 0.336 W and the energy efficiency at approximately 36,300 samples per joule. These results demonstrate that the proposed microarchitecture provides a resource-efficient solution for real-time neuromorphic edge inference, provided that the network size remains within the practical limits of time-multiplexed execution.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑