ARMOR:通过节点压缩缓解前端瓶颈以加速RTL仿真
ARMOR: Accelerating RTL Simulation by Mitigating the Front-End Bottleneck Using Node Compression
浏览论文内容
中文总结 AI 辅助
ARMOR是一款通过节点压缩缓解前端瓶颈的高效RTL仿真器,利用比特级数据并行性压缩节点,在CPU和AI加速器设计上分别实现1.6倍和2.7倍加速。
中文摘要 AI 辅助
RTL仿真在芯片设计中不可或缺,高性能仿真器通常会将RTL图中的每个节点降低为指令序列。尽管这种逐节点降低方式支持激进的编译器优化,但会大幅增加代码占用,严重超出指令缓存容量,进而引发前端瓶颈。我们的分析显示,超过50%的流水线停顿由CPU前端导致,成为最先进RTL仿真器的关键性能瓶颈。然而,在充分展开RTL图以获取优化收益的同时,又要减少代码占用以缓解前端瓶颈,这一目标仍极具挑战性。本文提出ARMOR,一款通过节点压缩缓解前端瓶颈的高效RTL仿真器。核心思路是利用展开后RTL图暴露的数据并行性,以及逐节点降低方式带来的比特空间利用不足,借助比特级数据并行性同时压缩多个节点,使单条指令序列可服务多个节点,而非每个节点对应一条指令序列。为实现高效的节点压缩,我们首先提出模块感知的同构子图识别方法,利用模块实例间的结构同构性,在子图层面系统识别压缩机会;接着提出对齐感知的密集打包策略,根据数据流依赖关系将节点分组为包,同时保留数据复用,辅以贪心合并策略提升比特空间利用率;最后实现统一的比特级并行方案,支持压缩节点的比特级并行执行。实验结果表明,与最先进的仿真器相比,ARMOR在CPU设计上实现1.6倍加速,在AI加速器上实现2.7倍加速。
英文摘要
RTL simulation is indispensable in chip design. High-performance simulators typically lower each node in the RTL graph into an instruction sequence. Although this per-node lowering enables aggressive compiler optimizations, it dramatically increases the code footprint, severely exceeding instruction cache capacity and causing front-end bottlenecks. Our profiling reveals that over 50% of pipeline stalls are caused by the CPU front-end, becoming a key performance bottleneck in state-of-the-art RTL simulators. However, reaping the optimization benefits of fully unrolling the RTL graph while simultaneously reducing the code footprint to mitigate front-end bottlenecks remains highly challenging. In this paper, we propose ARMOR, an efficient RTL simulator designed to alleviate the front-end bottleneck through node compression. The key idea is to exploit the data parallelism exposed by the unrolling RTL graph and the insufficient bit-space utilization revealed by per-node lowering, leveraging bit-level data parallelism to compress multiple nodes simultaneously, so that a single instruction sequence can serve multiple nodes instead of one per node. To achieve profitable node compression, we first propose a module-aware isomorphic subgraph identification method that leverages structural isomorphism across module instances to systematically identify compression opportunities at the subgraph level. We then propose an alignment-aware dense packing strategy that groups nodes into packs according to dataflow dependencies while preserving data reuse, complemented by greedy merging strategies to enhance bit-space utilization. Finally, we implement a unified bit-level parallelism scheme to support bit-level parallel execution of compressed nodes. Experimental results show that ARMOR achieves 1.6x speedup on CPU designs and 2.7x speedup on AI accelerators compared to state-of-the-art simulators.