发表机构
Intigia; DFKI(Intigia; 德国人工智能研究中心)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出一种基于FINN的NPU加速器,通过量化感知训练、模型精简和硬件优化,在PYNQ-Z1上实现满足实时性、能效和精度要求的车辆检测,达到35.66 FPS和0.594 mAP。
AI 中文摘要
本文介绍了在 PYNQ-Z1 板的资源受限的 Xilinx Zynq XC7Z020 器件上,用于实时车辆检测的神经处理单元(NPU)加速器的设计、优化、实现和板载验证。该工作遵循软硬件协同设计方法,结合了量化感知训练(QAT)、轻量级 YOLO 衍生检测器、Brevitas/QONNX 模型导出、FINN 数据流编译、Vivado 实现以及目标板上的物理基准测试。四个同时存在的工程需求定义了成功部署:吞吐量高于 30 帧/秒(FPS),能效高于 7 FPS/W,可编程逻辑(PL)硬件延迟低于 50 毫秒,以及 Pascal VOC 检测精度高于 0.55 mAP@0.5。设计空间包括 LP-YOLO 和 LP-YOLO Slim 变体、自定义 YOLOv3-tiny 参考、4 位和混合低比特量化、320×320 和 256×256 输入、手动和自动 FIFO 大小调整,以及从 100 到 200 MHz 的可编程逻辑时钟。最终的 LP-YOLO Slim 配置使用 256×256 输入、w2a4 量化和 142.86 MHz PL 时钟。在批处理大小为 100 时,它达到 35.66 FPS,功耗为 2.91 W,对应 12.25 FPS/W,而测得的 PL 延迟为 45.11 毫秒,VOC mAP@0.5 为 0.594。这是唯一一个所提供的测量结果同时满足所有四个需求的评估配置。结果表明,低比特 QAT、架构精简、FINN 折叠和 FIFO 优化以及适度的时钟缩放可以共同在小型 Zynq FPGA 上提供实用的实时检测器。
英文摘要
This paper presents the design, optimization, implementation, and on-board validation of a neural processing unit (NPU) accelerator for real-time vehicle detection on the resource-constrained Xilinx Zynq XC7Z020 device of the PYNQ-Z1 board. The work follows a hardware/software co-design methodology that combines quantization-aware training (QAT), lightweight YOLO-derived detectors, Brevitas/QONNX model export, FINN dataflow compilation, Vivado implementation, and physical benchmarking on the target board. Four simultaneous engineering requirements define successful deployment: throughput above 30 frames/s (FPS), energy efficiency above 7 FPS/W, programmable-logic (PL) hardware latency below 50 ms, and Pascal VOC detection accuracy above 0.55 mAP@0.5. The design space includes LP-YOLO and LP-YOLO Slim variants, a custom YOLOv3-tiny reference, 4-bit and mixed low-bit quantization, 320$\times$320 and 256$\times$256 inputs, manual and automatic FIFO sizing, and programmable-logic clocks from 100 to 200 MHz. The final LP-YOLO Slim configuration uses a 256$\times$256 input, w2a4 quantization, and a 142.86 MHz PL clock. With batch 100 it reaches 35.66 FPS at 2.91 W, corresponding to 12.25 FPS/W, while measured PL latency is 45.11 ms and VOC mAP@0.5 is 0.594. This is the only evaluated configuration for which the supplied measurements satisfy all four requirements simultaneously. The results show that low-bit QAT, architectural slimming, FINN folding and FIFO optimization, and moderate clock scaling can jointly provide a practical real-time detector on a small Zynq FPGA.