arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

FAME:一种基于FPGA的近似乘法器评估平台,具有模式引导的DNN重训练

FAME: An FPGA-Based Platform for Approximate Multipliers Evaluation with Pattern-Guided DNN Retraining

Rappy Saha, Nima Amirafshar, Jude Haris, Nima Taherinejad, José Cano

arXiv 2609.17730首次发表:更新:

发表机构

University of Glasgow; Heidelberg University(格拉斯哥大学; 海德堡大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

提出基于FPGA的FAME平台,直接硬件实现近似乘法器,加速DNN评估,并结合模式引导重训练恢复精度,在ImageNet上实现3.47倍加速和65.5%精度提升。

AI 中文摘要

近似乘法器可以减少深度神经网络(DNN)推理中的硬件面积和能耗;然而,它们会引入计算误差。由于评估时间过长,在多种DNN模型和大规模数据集上评估众多近似乘法器设计的准确性仍然具有挑战性。这一开销主要源于在CPU和GPU平台上使用查找表(LUT)对近似乘法器行为进行缓慢的仿真。此外,由此产生的精度下降必须仔细量化,并在必要时加以缓解(例如通过重训练),这进一步增加了整体评估成本。为了解决这些挑战,我们提出了FAME,一个基于FPGA的近似乘法器评估平台。该平台利用现场可编程门阵列(FPGA)的可重构逻辑直接在硬件中实现近似乘法器,消除了在CPU/GPU平台上基于LUT的仿真需求,从而能够高效地进行DNN推理,同时显著减少大规模数据集上的评估时间。此外,我们引入了一种模式引导的DNN重训练技术,以减轻由近似乘法器引起的精度下降。具体而言,重训练由乘法器特定的模式引导,以有效恢复潜在的精度损失。我们使用两种DNN模型ResNet-18和MobileNetV2,在ImageNet数据集上对27个近似乘法器评估了FAME。在推理过程中,与先前基于LUT的仿真方法相比,我们的方法在近似乘法器评估中实现了高达3.47倍的加速。此外,所提出的重训练技术相对于现有的重训练方法,在评估的乘法器上将精度提高了高达65.5%。代码公开于:此https URL。

英文摘要

Approximate multipliers can reduce hardware area and energy consumption in Deep Neural Network (DNN) inference; however, they introduce computational errors. Assessing the accuracy of numerous approximate multiplier designs across diverse DNN models and large-scale datasets remains challenging due to prohibitive evaluation times. This overhead primarily stems from the slow emulation of approximate multiplier behavior using look-up tables (LUTs) on CPU and GPU platforms. Moreover, the resulting accuracy degradation must be carefully quantified and, if necessary, mitigated (e.g., through retraining), further increasing the overall evaluation cost. To address these challenges, we propose FAME, an FPGA-based platform for evaluating approximate multipliers. The platform exploits the reconfigurable logic of Field-Programmable Gate Arrays (FPGAs) to implement approximate multipliers directly in hardware, eliminating the need for LUT-based emulation on CPU/GPU platforms and thereby enabling efficient DNN inference while significantly reducing evaluation time on large datasets. Furthermore, we introduce a pattern-guided DNN retraining technique to mitigate accuracy degradation induced by approximate multipliers. Specifically, retraining is guided by multiplier-specific patterns to effectively recover potential accuracy losses. We evaluate FAME using two DNN models, ResNet-18 and MobileNetV2, on the ImageNet dataset across 27 approximate multipliers. During inference, our approach achieves up to a 3.47x speedup in approximate multiplier evaluation compared to prior LUT-based emulation methods. Furthermore, the proposed retraining technique improves accuracy by up to 65.5% over existing retraining approaches for the evaluated multipliers. The code is publicly available at: https://github.com/gicLAB/FAME

CommentsAccepted at the 38th IEEE/SBC International Symposium on Computer Architecture and High Performance Computing (SBAC-PAD) 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑