MCHA:一种面向并行-串行计算的以内存为中心的分层架构
MCHA: A Memory-Centric Hierarchical Architecture for Parallel-Sequential Computing
浏览论文内容
中文总结 AI 辅助
针对多智能体强化学习等并行-串行计算工作负载的内存瓶颈,提出以内存为中心的分层架构MCHA,通过分层通信策略和并行-串行编程模型实现大幅性能提升并开源。
中文摘要 AI 辅助
多智能体强化学习(MARL)、大规模神经形态计算和概率图模型等新兴工作负载固有地呈现出并行-串行计算模式。这些任务需要大量并行性来实现高吞吐量,但却因集中在主存的不规则数据访问模式而面临严重瓶颈。因此,传统架构在执行这些工作负载时存在根本性限制,主要表现为全局缓冲区饱和和内存受限瓶颈。为应对这些挑战,我们提出了以内存为中心的分层架构(Memory-Centric Hierarchical Architecture, MCHA),这是一种专为并行-串行执行定制的可重构硬件解决方案。MCHA利用分层通信策略促进分布式核间数据路由,从而显著降低全局内存的带宽负担。除硬件外,MCHA还引入了一种新颖的并行-串行编程模型,该模型利用事件驱动的条件触发器,可在执行流水线中有效隐藏数据传输延迟。我们针对多种并行-串行任务对MCHA进行了基准测试,包括MARL、电机变量控制和马尔可夫随机场。通过我们开源的周期精确模拟器验证,在MARL工作负载上,MCHA相比NVIDIA A100 GPU实现了153.06倍至2456.96倍的性能加速,同时在其他应用领域保持了强大的编程灵活性。此外,该架构将主存访问从96%降至5.44%。在28nm工艺下综合时,MCHA实现的面积为2.92mm²,在200MHz频率下功耗为115.36mW。MCHA已在该httpsURL开源。
英文摘要
Emerging workloads, such as Multi-Agent Reinforcement Learning (MARL), large-scale neuromorphic computing, and probabilistic graphical models, intrinsically exhibit parallel-sequential computing patterns. While these tasks demand massive parallelism to achieve high throughput, they are severely bottlenecked by irregular data access patterns centralized to main memory. Consequently, conventional architectures face fundamental limitations when executing these workloads, primarily manifesting as global buffer saturation and memory-bound bottlenecks. To address these challenges, we propose the Memory-Centric Hierarchical Architecture (MCHA), a reconfigurable hardware solution tailored for parallel-sequential execution. MCHA leverages a hierarchical communication strategy that facilitates distributed, inter-core data routing, thereby significantly reducing the bandwidth burden on the global memory. Complementing the hardware, MCHA introduces a novel parallel-sequential programming model that utilizes event-driven conditional triggers to effectively hide data transmission latency within the execution pipeline. We benchmark MCHA against a diverse suite of parallel-sequential tasks, including MARL, motor variable control, and Markov random fields. Validated through our open-source, cycle-accurate simulator, MCHA demonstrates performance speedups ranging from 153.06$\times$ to 2456.96$\times$ over NVIDIA A100 GPUs on MARL workloads, while maintaining robust programming flexibility across other application domains. Furthermore, the architecture successfully reduces main memory access from 96% to 5.44%. When synthesized in a 28 nm process, the MCHA implementation occupies an area footprint of 2.92mm$^2$ and consumes 115.36 mW of power at 200 MHz. MCHA is open-sourced at https://github.com/carabdis/MCHA.