无需注意力机制:FPGA上现代YOLO中基于DPU的注意力近似方法
No Attention, No Problem: DPU-Aware Attention Approximation in Modern YOLO on FPGA
浏览论文内容
中文总结 AI 辅助
研究针对FPGA上基于注意力的YOLO变体,提出DPU感知架构,评估YOLOv26和YOLOv11,替换函数、近似注意力机制,在多数据集训练评估,在多种DPU配置下测试,得出YOLOv26n等吞吐量高、精度降5%、功耗降约3倍的结果。
中文摘要 AI 辅助
基于边缘的人工智能加速技术在实时目标检测方面取得了进展。边缘设备上的目标检测需要在准确性、速度和功率效率之间取得平衡。本文针对部署在AMD FPGA上基于注意力机制的YOLO变体提出了一种定制的深度学习处理器单元(DPU)感知架构。具体评估了YOLOv26和YOLOv11在标准和定向目标检测任务中的性能,对激活函数等进行了替换和近似处理。在多个数据集上训练和评估模型,并在所有八个DPU配置下进行基准测试。结果表明,YOLOv26n和YOLOv26n-OBB分别在标准和定向检测中实现了最高的端到端吞吐量,量化导致平均精度绝对降低5%,但功耗降低约3倍。
英文摘要
Edge-based Artificial Intelligence (AI) acceleration has recently improved progress in real-time object detection. Object detection on edge devices requires a balance between accuracy, speed, and power efficiency. This paper proposes a customized Deep Learning Processor Unit (DPU)-aware architecture for attention-based YOLO variants deployed on AMD FPGAs. Specifically, we evaluate and benchmark YOLOv26 and YOLOv11, two modern attention-based YOLO variants, on the Xilinx ZCU104 across both standard and oriented object detection tasks. We replace unsupported activation functions, substitute split operations with 1x1 convolutions, and approximate the spatial attention mechanism in a DPU-compatible way. All models are then trained and evaluated across six benchmark datasets such as COCO, Pascal VOC, KITTI, DOTA, DIOR-R, and an in-house human presence dataset, and benchmarked across all eight DPU configurations (B512 to B4096) in terms of mAP, FPS, latency, power, and resource utilization. Notably, YOLOv26n and YOLOv26n-OBB deliver the highest end-to-end throughput at 34.05 and 29.55 FPS for standard and oriented detection, respectively, with an average of 5% absolute reduction in accuracy due to quantization while achieving up to approximately 3x lower power consumption compared with the state of the art.