arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

面向边缘视觉的基于优先分配空闲任务的低功耗稀疏卷积加速器

A Low-Power Sparse Convolution Accelerator with Idle-First-Task-Assignment for Edge Vision

Jingyue Zhuge, Johannes Partzsch, Christian Mayr

arXiv 2607.26835首次发表:更新:

AI 中文总结

本文针对边缘视觉监控系统的资源约束,提出基于IFTA调度的低功耗稀疏卷积加速器,采用16 nm工艺实现,在ImageNet上对稀疏VGG16、MobileNetV2的加速效果优于传统密集及现有稀疏加速器。

AI 中文摘要

近年来,智能畜牧等应用的边缘视觉监控系统面临严格的三方约束:在极有限的传输带宽和严格的功耗预算下保持输入分辨率。传统密集卷积神经网络(CNNs)无法满足此类受限物联网节点的资源限制。为应对这一挑战,本文提出一种面向边缘设备的低功耗稀疏卷积加速器,采用16 nm工艺流片并验证。首先,该加速器在数据传输和计算中采用基于位图的格式进行压缩,有效降低内存和带宽开销。其次,为缓解稀疏计算中的负载不平衡,提出优先分配空闲任务(Idle-First-Task-Assignment,IFTA)动态调度策略,显著降低处理单元(PE)空闲时间并提升乘法器利用率。此外,设计专用数据流以支持和加速轻量化网络中广泛使用的深度可分离卷积(DWConv)。实验结果表明,该芯片核心面积仅为0.5 mm²,功耗低至12~16 mW;在ImageNet上,针对稀疏VGG16和MobileNetV2,该加速器相比传统密集加速器分别实现6.5倍和2.8倍的加速,且相比现有稀疏加速器也取得显著性能提升。

英文摘要

In recent years, edge-vision monitoring systems for applications such as smart animal husbandry have faced strict tripartite constraints: maintaining input resolution under extremely limited transmission bandwidth and strict power budgets. Conventional dense convolutional neural networks (CNNs) cannot satisfy the resource limits of such constrained IoT nodes. To address this challenge, this paper presents a low-power sparse convolution accelerator for edge devices, fabricated and validated in a 16 nm process. First, the accelerator adopts a bitmap-based format for compression in both data transmission and computation, effectively reducing memory and bandwidth overhead. Second, to mitigate load imbalance in sparse computation, an Idle-First-Task-Assignment (IFTA) dynamic scheduling strategy is proposed, significantly reducing processing-element (PE) idle time and improving multiplier utilization. In addition, a dedicated dataflow is designed to support and accelerate depthwise separable convolution (DWConv), which is widely used in lightweight networks. Experimental results show that the chip occupies only 0.5~mm$^2$ core area and consumes as little as 12--16~mW. On ImageNet, for sparse VGG16 and MobileNetV2, the proposed accelerator achieves 6.5$\times$ and 2.8$\times$ speedups, respectively, over traditional dense accelerators, and also delivers significant performance gains over the existing sparse accelerator.

Comments5 pages, 5 figures, Accepted at the 2026 IEEE 8th International Conference on Artificial Intelligence Circuits and Systems (AICAS 2026)

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑