arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

CMD:一种具有基于簇的分布式存储器设计的集成CGRA框架

CMD: An Integrated CGRA Framework with Cluster-Based Distributed Memory Design

Shangkun Li, Cheng Tan, Zeyu Li, Jinming Ge, Jiawei Liang, Hao Yang, Linfeng Du, Jiang Xu, Wei Zhang

arXiv 2609.05982首次发表:更新:

发表机构

The Hong Kong University of Science and Technology; Google; Arizona State University; The George Washington University; The Hong Kong University of Science and Technology (Guangzhou)(香港科技大学; 谷歌; 亚利桑那州立大学; 乔治华盛顿大学; 香港科技大学(广州))

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

提出CMD框架,通过基于簇的分布式存储器和协同设计的编译工具链,解决CGRA存储瓶颈,实现1.39倍加速并减少面积。

AI 中文摘要

粗粒度可重构阵列(CGRA)是跨各种应用领域实现高能效和可重构性的有前景的解决方案,但其性能常常受限于僵化的存储器架构,这种架构限制了能够访问数据存储器的瓦片(tile)的数量和位置。这给具有密集存储器访问的内核带来了显著的瓶颈。为了解决这一问题,我们提出了CMD,一种集成CGRA框架,其特点是具有基于簇的分布式存储器设计,并配有协同设计的编译工具链。该编译器包含一种新颖的存储器感知映射器,以及一种设计空间探索(DSE)机制,该机制为特定内核确定最优的存储器架构设计。实验结果表明,经过DSE后的CMD CGRA相比传统CGRA实现了平均$1.39\ imes$的加速,同时将总面积平均降低到传统CGRA的$0.912\ imes$。

英文摘要

Coarse-Grained Reconfigurable Arrays (CGRAs) are a promising solution for achieving high energy efficiency and reconfigurability across various application domains, but their performance is often crippled by rigid memory architectures that limit the number and location of tiles that can access data memory. This creates a significant bottleneck for kernels with intensive memory accesses. To address this, we propose CMD, an integrated CGRA framework featuring cluster-based distributed memory design with a co-designed compilation toolchain. The compiler includes a novel memory-aware mapper and a design space exploration (DSE) mechanism that identifies the optimal memory architecture design for specific kernels. Experimental results show that our post-DSE CMD CGRAs achieve an average speedup of $1.39\times$ over a conventional CGRA while simultaneously reducing the total area to an average of $0.912\times$ of the conventional CGRA.

CommentsAccepted by ICCD 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑