arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

STELLA:一种面向多级流水线应用的16nm时空弹性低延迟粗粒度可重构阵列

STELLA: A 16nm Spatio-Temporal Elastic Low-Latency CGRA for Multi-Stage Pipelined Applications

Jun Yin, Chao Fang, Ryan Antonio, Xiaoling Yi, Yunhao Deng, Fanchen Kong, Marian Verhelst

arXiv 2609.39703首次发表:更新:

发表机构

ESAT-MICAS, KU Leuven(荷语鲁汶大学 ESAT-MICAS 研究中心)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

STELLA提出一种16nm时空弹性低延迟CGRA,通过快速配置、硬件循环控制和深度流水线弹性结构,在850MHz下实现110 GOPS/mm2,内核吞吐量较基线提升4.84-7.14倍。

AI 中文摘要

新兴的非矩阵机器学习内核,如LayerNorm、GeLu、FFT或循环卷积,需要低延迟、高能效的空间加速器,而非仅以MatMul为中心的阵列。STELLA提出了一种时空弹性的16nm粗粒度可重构阵列(CGRA),具有快速配置路径、每处理单元(PE)硬件循环控制,以及具备时空数据复用的低延迟、深度流水线弹性结构。STELLA在850 MHz下达到高达110 GOPS/mm2的性能,并将有效内核吞吐量较基线CGRA提升了4.84-7.14倍。

英文摘要

Emerging non-matrix ML kernels, such as LayerNorm, GeLu, FFT or circular convolutions, demand low-latency, energy-efficient spatial accelerators beyond MatMul-centric arrays. STELLA presents a spatio-temporal elastic 16 nm coarse-grained reconfigurable array (CGRA) with a rapid configuration path, per-PE hardware loop control, and a low-latency, deeply pipelined elastic fabric with spatio-temporal data reuse. STELLA reaches up to 110 GOPS/mm2 at 850 MHz, and improves effective kernel throughput by 4.84-7.14x over baseline CGRAs.

Journal ref2026 IEEE Custom Integrated Circuits Conference (CICC), Seattle, WA, USA. IEEE, 2026

DOI:10.1109/CICC65509.2026.11509531

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑