AI 中文总结
本综述针对FPGA的深度学习部署,建立压缩与硬件加速的分类体系,分析25项案例研究,发现领域表征缺口,明确六项开放挑战并给出具体下一步方向。
AI 中文摘要
将深度神经网络部署在现场可编程门阵列(Field-Programmable Gate Arrays,FPGAs)上,需要同时考虑模型压缩与硬件加速,但现有最全面的跨平台研究(Deng等人,2020)仅在宽泛的定性权衡层面对比了针对CPU、GPU、FPGA和ASIC目标的压缩技术,未涉及具体的FPGA资源影响。本综述将范围限定为FPGA,把2015至2026年间的25项压缩-硬件协同设计案例研究,按各策略主要重塑的FPGA资源分为五类:消除数字信号处理器(DSP)、DSP再利用/混合精度、利用稀疏性、驱动内存层次结构、工具链/部署级。通过在统一维度(压缩率、精度变化、吞吐量、能效、DSP/LUT/BRAM利用率)上标准化这些案例研究,得出核心定量发现:25项研究中仅1项报告了针对单一公共基准的压缩率与精度变化,仅2项报告了针对公共GPU基准的标准化能效,暴露出该领域存在的表征缺口,而FINN、HLS4ML、Vitis AI或DNNWeaver等工具链均无法单独解决此问题。基于该分类体系与元分析,本文明确了六项开放挑战:工具链碎片化、精度-能效表征、自动化混合精度优化、稀疏计算可靠性、持久内存瓶颈、基于FPGA的训练,并为每项挑战提出了基于扩展现有引用技术的具体下一步,而非泛泛的未来工作呼吁。
英文摘要
Deploying deep neural networks on Field-Programmable Gate Arrays (FPGAs) requires joint reasoning about model compression and hardware acceleration, however the most comprehensive existing cross-platform treatment of this space, Deng et al.~\cite{deng2020model}, compared compression techniques against CPU, GPU, FPGA, and ASIC targets at the level of broad, qualitative trade-offs, and not specific FPGA resource consequences. This survey instead restricted the scope to FPGAs alone and organized 25 compression-hardware co-design case studies (2015--2026) into a five-category taxonomy defined by which FPGA resources each strategy primarily reshapes: DSP-eliminating, DSP-repurposing/mixed-precision, sparsity-exploiting, memory-hierarchy-driven, and toolchain/deployment-level. Normalizing these case studies along a common set of dimensions (compression ratio, accuracy change, throughput, energy efficiency, and DSP/LUT/BRAM utilization) surfaces a central, quantitative finding; of the 25 reviewed works, only \emph{one} reported a compression ratio and accuracy change measured against a single common baseline, and only \emph{two} reported energy efficiency normalized against a common GPU baseline, exposing a field-wide characterization gap that no individual toolchain (FINN, HLS4ML, Vitis AI, or DNNWeaver) resolves on its own. Building on this taxonomy and meta-analysis, we formalize six open challenges: toolchain fragmentation, accuracy--efficiency characterization, automated mixed-precision optimization, sparse computation reliability, persistent memory bottlenecks, and FPGA-based training. Each is paired with a concrete next step grounded in extending an existing, cited technique, not a general call for future work.