发表机构
HKUST; ETH Zurich; HKUST (GZ)(香港科技大学; 苏黎世联邦理工学院; 香港科技大学(广州))
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究旨在实现DNN纳秒级推理延迟,针对传统FPGA加速器局限,提出FPGN框架。通过硬件对齐可微公式、结构化拓扑和延迟驱动编译器,弥合差距。实验表明其相比其他加速器延迟降低205倍,LUT效率高30倍,且推理精度有竞争力。
AI 中文摘要
实现深度神经网络(DNN)的纳秒级推理延迟已成为对延迟敏感应用的主要架构关注点。虽然现场可编程门阵列(FPGA)为低延迟推理提供了有前景的基础,但传统FPGA加速器仍以算术为中心,主要将查找表(LUT)用作数值运算符和外围逻辑的构建块。相比之下,最近的原生LUT神经网络将LUT视为可学习神经元,显示出利用其内在逻辑表达能力的有前景的理论潜力。然而,现有方法主要局限于算法优化,未能将这种理论潜力转化为高性能FPGA加速器。具体而言,它们的可微公式与FPGA LUT原语不完全匹配,其物理上无感知的拓扑结构损害了可路由性和时序收敛,并且缺乏自动优化流程阻碍了系统设计空间探索(DSE)和高效硬件实现。在本文中,我们提出了FPGN,这是一个端到端的物理感知框架,弥合了原生LUT学习与延迟优化FPGA实现之间的差距。FPGN通过(i)用于训练FPGA原生LUT神经元的硬件对齐可微公式,(ii)具有流硬件架构的结构化原生LUT拓扑结构以改善路由局部性和时序收敛,以及(iii)利用高保真分析结果质量模型自动进行DSE和硬件生成的延迟驱动编译器来应对这些挑战。实验表明,与基于FPGA的代表性BNN加速器相比,FPGN实现了高达205倍的延迟降低,并且比先前的可微原生LUT网络具有高达30倍的LUT效率,同时保持有竞争力的推理精度。
英文摘要
Achieving nanosecond-scale inference latency for deep neural networks (DNNs) has become a primary architectural concern for latency-critical applications. While Field-Programmable Gate Arrays (FPGAs) offer a promising substrate for low-latency inference, conventional FPGA accelerators remain arithmetic-centric, using LUTs primarily as building blocks for numerical operators and peripheral logic. In contrast, recent LUT-native neural networks treat LUTs as learnable neurons, revealing promising theoretical potential to exploit their intrinsic logic expressivity. However, existing methods are largely confined to algorithmic optimizations, failing to translate this theoretical potential into high-performance FPGA accelerators. Specifically, their differentiable formulations do not faithfully match FPGA LUT primitives, their physically-unaware topologies compromise routability and timing closure, and their lack of automated optimization flow hinders systematic design space exploration (DSE) and efficient hardware implementation. In this paper, we propose FPGN, an end-to-end physically-aware framework that closes the gap between LUT-native learning and latency-optimized FPGA implementation. FPGN addresses these challenges through (i) a hardware-aligned differentiable formulation for training FPGA-native LUT neurons, (ii) a structured LUT-native topology with a streaming hardware architecture to improve routing locality and timing closure, and (iii) a latency-driven compiler that leverages high-fidelity analytical Quality of Results models to automate DSE and hardware generation. Experiments show that FPGN achieves up to 205$\times$ latency reduction compared to representative FPGA-based BNN accelerators and up to 30$\times$ higher LUT efficiency than prior differentiable LUT-native networks, while maintaining competitive inference accuracy.