发表机构
University College Dublin; Virginia Tech; Queen’s University Belfast(都柏林大学学院; 弗吉尼亚理工大学; 贝尔法斯特女王大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出ZTA-Q,一个基于RISC-V的开源平台,用于精确部署TensorFlow Lite INT8 CNN模型,通过可配置后处理数据通路研究电路级近似对精度的影响,在FPGA上以83.3 MHz运行,将精度损失控制在0.25个百分点内。
AI 中文摘要
低精度推理在边缘AI中被广泛采用,以降低计算成本和内存占用。然而,现有的开源加速器平台对遵循标准TensorFlow Lite整数推理方案的CNN提供了有限的全端到端支持。本文提出了ZTA-Q,一个基于RISC-V的开源平台,能够实现TensorFlow Lite INT8模型的精确部署。除了扩展算子支持外,ZTA-Q还提供了一个可配置的后处理数据通路,用于研究电路级近似(包括降低乘法器精度、共享移位缩放和简化舍入)如何影响模型精度。所提出的系统在Digilent Arty A7-100T FPGA上实现,并以83.3 MHz的频率运行。对代表性CNN模型的评估表明,在LUT、寄存器和DSP开销分别为26.3%、12.6%和150%的情况下,ZTA-Q将top-1和top-5精度的下降限制在0.25个百分点以内。
英文摘要
Low-precision inference is widely adopted in edge AI to reduce computational cost and memory footprint. However, existing open-source accelerator platforms provide limited end-to-end support for CNNs following the standard TensorFlow Lite integer inference scheme. This paper presents ZTA-Q, an open-source RISC-V-based platform that enables accurate deployment of TensorFlow Lite INT8 models. In addition to extending operator support, ZTA-Q provides a configurable post-processing datapath for studying how circuit-level approximations, including reduced multiplier precision, shared shift scaling, and simplified rounding, affect model accuracy. The proposed system is implemented on a Digilent Arty A7-100T FPGA and operates at 83.3 MHz. Evaluations on representative CNN models show that with LUT, register, and DSP overheads of 26.3%, 12.6%, and 150%, respectively, ZTA-Q limits the degradation in both top-1 and top-5 accuracy to within 0.25 percentage points.
CommentsAccepted for publication at the 2026 IEEE 33rd International Conference on Electronics, Circuits and Systems (ICECS)