发表机构
ESAT-MICAS, KU Leuven(荷语鲁汶大学 ESAT-MICAS 研究中心)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
STELLA提出一种16nm时空弹性低延迟CGRA,通过快速配置、硬件循环控制和深度流水线弹性结构,在850MHz下实现110 GOPS/mm2,内核吞吐量较基线提升4.84-7.14倍。
AI 中文摘要
新兴的非矩阵机器学习内核,如LayerNorm、GeLu、FFT或循环卷积,需要低延迟、高能效的空间加速器,而非仅以MatMul为中心的阵列。STELLA提出了一种时空弹性的16nm粗粒度可重构阵列(CGRA),具有快速配置路径、每处理单元(PE)硬件循环控制,以及具备时空数据复用的低延迟、深度流水线弹性结构。STELLA在850 MHz下达到高达110 GOPS/mm2的性能,并将有效内核吞吐量较基线CGRA提升了4.84-7.14倍。
英文摘要
Emerging non-matrix ML kernels, such as LayerNorm, GeLu, FFT or circular convolutions, demand low-latency, energy-efficient spatial accelerators beyond MatMul-centric arrays. STELLA presents a spatio-temporal elastic 16 nm coarse-grained reconfigurable array (CGRA) with a rapid configuration path, per-PE hardware loop control, and a low-latency, deeply pipelined elastic fabric with spatio-temporal data reuse. STELLA reaches up to 110 GOPS/mm2 at 850 MHz, and improves effective kernel throughput by 4.84-7.14x over baseline CGRAs.
Journal ref2026 IEEE Custom Integrated Circuits Conference (CICC), Seattle, WA, USA. IEEE, 2026
DOI:10.1109/CICC65509.2026.11509531