发表机构
The Hong Kong University of Science and Technology - Guangzhou Campus; Stevens Institute of Technology(香港科技大学广州校区; 史蒂文斯理工学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出首个GPU加速的门级时域功耗分析框架,通过事件密度感知分区和内核融合,实现高达37.63倍加速,同时保持高精度。
AI 中文摘要
功耗分析在现代芯片设计流程中至关重要。特别是,时域功耗分析能够提供细粒度的功耗消耗信息,有助于诊断功耗问题并据此指导功耗优化。然而,对现代大规模电路进行时域功耗分析可能需要数十小时,这极大地减慢了功耗优化流程。在本文中,我们提出了首个GPU加速的门级时域功耗分析框架。我们提出了一种新颖的数据结构,以实现高效的状态相关功耗检索。为了适应门之间不平衡的事件分布,我们提出了一种事件密度感知的分区策略,该策略根据门事件密度分配GPU线程。最后,我们将功耗计算融合到单次内核调用中,以减少单独内核中的冗余工作。实验结果表明,我们提出的框架实现了高精度,同时与多线程的Synopsys PrimeTime PX相比,提供了高达37.63倍的端到端加速。
英文摘要
Power analysis is crucial in modern chip design flow. Particularly, time-based power analysis can provide fine-grained power consumption information to facilitate the diagnosis of power issues and guide power optimization accordingly. However, it may take tens of hours to conduct time-based power analysis on modern large-scale circuits, which greatly slows down the power optimization flow. In this paper, we present the first GPU-accelerated gate-level time-based power analysis framework. We propose a novel data structure to enable efficient state-dependent power retrieval. To accommodate the imbalanced event distribution across gates, we propose an event-density-aware partitioning strategy that allocates GPU threads based on gate event density. Finally, we fuse the power computation into a single kernel invocation to reduce redundant work in separate kernels. Experimental results show that our proposed framework achieves high accuracy while delivering up to 37.63x end-to-end speedup compared to multi-threaded Synopsys PrimeTime PX.
CommentsAccepted for publication in ACM Transactions on Architecture and Code Optimization (TACO)