arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

一种具有可配置MAC和动态截断的420 GOPS/W CGRA

A 420 GOPS/W CGRA with a Configurable MAC and Dynamic Truncation

Yi Sheng Chong, Rakshith Harish, Rajesh Chandrasekhara Panicker, Vishnu P. Nambiar, Anh Tuan Do

arXiv 2609.16600首次发表:更新:

发表机构

Institute of Microelectronics, Agency for Science, Technology and Research (A*STAR); Department of Electrical and Computer Engineering, National University of Singapore (NUS)(微电子研究所,科技研究局; 新加坡国立大学电气与计算机工程系)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对边缘设备实时负载需求,提出一种单周期可配置MAC的CGRA,采用40nm工艺,能效达420.6 GOPS/W,较先进技术提升1.4倍。

AI 中文摘要

边缘设备对高效且灵活的处理能力有强烈需求,以处理动态实时工作负载。粗粒度可重构架构(CGRA)因其既具有通用处理器的灵活性,又具有接近领域专用加速器的高效率,而成为边缘设备中合适的加速器候选。然而,典型的CGRA执行一次乘累加(MAC)操作需要两个周期,而神经网络推理和信号处理等工作负载涉及大量MAC操作,导致CGRA处理时间较长。本工作提出了一种CGRA,其处理单元(PE)中具有可配置的MAC单元,通过使用相同的乘法器和加法器,可以在单个周期内执行加法(ADD)、乘法(MUL)或MAC操作。MAC结果的读出精度可通过截断块进行调整。所提出的CGRA采用40nm CMOS工艺实现。在0.6V电源电压和21MHz频率下工作时,其能效达到420.6 GOPS/W,是现有最先进技术的1.4倍。

英文摘要

Edge devices demand for highly efficient yet flexible processing capability to handle dynamic real-time workloads. Coarse grain reconfigurable architecture (CGRA) emerges as a suitable accelerator candidate in edge devices, because they are as flexible as general purpose processors and offer high efficiency close to that of domain specific accelerators. However, a typical CGRA requires two cycles for a multiply-and-accumulate (MAC) operation, and workloads such as neural network inference and signal processing involve many MAC operations, resulting in long CGRA processing time. This work proposes a CGRA that has configurable MAC units in the processing elements (PEs) that can perform an addition (ADD) or multiplication (MUL) or a MAC by using the same multiplier and adder, in a single cycle. The readout precision of MAC result can be adjusted by a truncation block. The proposed CGRA is implemented with 40nm CMOS technology. It attains an energy efficiency of 420.6GOPS/W operating at supply of 0.6V and frequency of 21MHz, which is 1.4 times higher than the state-of-the-art.

Comments5 pages, 8 figures, accepted by 2024 IEEE International Symposium on Circuits and Systems (ISCAS)

DOI:10.1109/ISCAS58744.2024.10558192

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑