arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.30954cs.AR

开源MSP430内核上的时钟门控插入策略:可复现的PPA研究与门级仿真注意事项

Clock-Gating Insertion Strategies on an Open-Source MSP430 Core: A Reproducible PPA Study and a Gate-Level Simulation Caveat

  • University of California, Riverside(加州大学河滨分校)
  • Texas A&M University–Corpus Christi(德克萨斯农工大学科珀斯克里斯蒂分校)
  • University of Missouri(密苏里大学)
  • University of Texas at Arlington(德克萨斯大学阿灵顿分校)
  • Boston University(波士顿大学)

机构由 AI 辅助整理,请以论文原文为准。

Xingran Huang, Qiming Guo, Jinwen Tang, Wenqi Jia, Dongzheng Wang

AI总结:

本文以openMSP430内核为研究对象,对比RTL行为时钟门与工具插入的ICG单元的PPA表现,发现二者门级仿真存在差异,建议开源低功耗设计优先选用ICG单元。

AI中文摘要:

时钟门控是削减动态功耗的标准技术,可在寄存器传输级(RTL)以手写行为时钟门的形式引入,或在综合过程中自动插入集成时钟门(ICG)单元,二者常被视为可互换。本文在采用32nm标准单元库综合的真实开源16位微控制器内核openMSP430上证明,二者实际并不等效:基于锁存器的RTL行为门控在理想RTL仿真中功能正确(10个自检查测试用例均通过,与未门控基线一致),但在门级仿真中失效——门控乘法器结果从未被捕获且读数为零,而工具插入的ICG单元门级仿真全部通过(10/10)。我们将失效根因定位为迟滞的锁存器+与门门控时钟引入的保持竞争,且该问题在8种仿真配置(含完整标准延迟格式SDF反标注)中均存在,并非仿真器设置的人为误差。随后我们针对未门控基线,在4种工作负载和3种工艺角(ss/tt/ff)下,量化了三种门控强度的功耗/面积/时序(PPA)影响:RTL行为门控(Opt1)、综合ICG(Opt2)及二者结合(Opt3)。该优势具有工艺角鲁棒性:ICG(Opt2)在所有工艺角下可削减74-81%的动态功耗和25-30%的总功耗。我们还表明,在这种漏电流主导的32nm工艺下,总功耗优势来自门控带来的面积/漏电流削减(漏电流降低24-30%),而非占活跃模式能耗主导的大幅动态功耗节省。我们对开源内核低功耗设计的建议是优先选用工具插入的ICG单元而非手写行为时钟门。完整流程(Design Compiler综合、PrimeTime PX功耗分析及自检查验证)作为开源工件发布。

英文摘要:

Clock gating, the standard technique for cutting dynamic power, is introduced either as hand-written behavioral clock gates at the register-transfer level (RTL) or as integrated clock-gating (ICG) cells inserted automatically during synthesis; the two are widely treated as interchangeable. In this paper we show, on a real open-source 16-bit microcontroller core (openMSP430) synthesized with a 32 nm standard-cell library, that they are not equivalent in practice: behavioral latch-based RTL gating is functionally correct in ideal RTL simulation (10/10 self-checking testcases, identical to the ungated baseline) yet fails at gate level: the gated multiplier result is never captured and reads zero, while tool-inserted ICG cells pass gate-level simulation cleanly (10/10). We root-cause the failure to a hold race introduced by the late latch+AND gated clock, and show it persists across eight simulation configurations including full Standard Delay Format (SDF) back-annotation, not a simulator-setting artifact. We then quantify the power/area/timing (PPA) impact of three gating strengths: RTL behavioral (Opt1), synthesis ICG (Opt2), and both (Opt3), against the ungated baseline, across four workloads and three process corners (ss/tt/ff). The benefit is corner-robust: ICG (Opt2) cuts dynamic power by 74-81% and total power by 25-30% at every corner. We also show that in this leakage-dominated 32 nm regime the total-power win comes from the area/leakage reduction that gating brings (leakage -24 to -30%), not from the large dynamic saving, which instead dominates active-mode energy. Our recommendation for low-power design on open-source cores is to prefer tool-inserted ICG cells over hand-written behavioral clock gates. The full flow (Design Compiler synthesis, PrimeTime PX power, and self-checking verification) is released as an open artifact.

补充信息

↑