arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.19837cs.AR

基于FFT的卷积神经网络与自适应浮点量化的节能视觉检测

Energy-Efficient Visual Inspection with FFT-Based CNNs and Adaptive Floating-Point Quantization

Lukas Krupp, Marco Groß, Michael Graichen, Kim Ulrich, Norbert Wehn

AI总结:

该研究针对工业CPU-FPGA平台,结合FFT卷积与自适应FP8量化优化LeNet-5加速器,实现故障检测准确率提升与能效增强。

AI中文摘要:

本文研究工业CPU-FPGA平台上基于FFT的卷积神经网络(CNN)推理的低精度浮点算术。将基于FFT的卷积与自适应训练后FP8量化结合,评估两种面向FPGA的优化方法:FFT内的渐进式偏差调整(PBA)和CNN各层的指数偏差选择。这些方法在采用串行基-2²SDF FFT模块的LeNet-5加速器中实现,并在工业故障检测数据集上评估。结果显示,权重缩放优于PBA,而层偏差优化可在不修改数据通路宽度的情况下将准确率从80.33%提升至84.13%;与仅CPU推理相比,FPGA的能效提升约2.5倍。

英文摘要:

This paper investigates reduced-precision floating-point arithmetic for FFT-based CNN inference on an industrial CPU-FPGA platform. We combine FFT-based convolution with adaptive post-training FP8 quantization and evaluate two FPGA-oriented optimization methods: progressive bias adjustment (PBA) within the FFT and layer-wise exponent-bias selection across the CNN. The methods are implemented in a LeNet-5 accelerator using serial radix-$2^2$ SDF FFT modules and evaluated on an industrial fault detection dataset. Results show that weight scaling outperforms PBA, while layer-wise bias optimization increases the accuracy from 80.33% to 84.13% without modifying the datapath width. Compared with CPU-only inference, the FPGA achieves approximately 2.5$\times$ higher energy efficiency.

补充信息

↑