arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

动态松弛感知时钟用于近阈值张量处理单元(TPU)

Dynamic Slack-Aware Clocking for Near-Threshold Tensor Processing Units (TPUs)

Muhammad Usman Nadeem, Sanghamitra Roy, Koushik Chakraborty

arXiv 2609.26644首次发表:更新:

发表机构

Utah State University(犹他州立大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

提出动态松弛感知时钟(DSAC),通过预测MAC操作延迟并分级调整时序,在近阈值TPU上实现最高1.55倍能效提升,精度损失仅1%。

AI 中文摘要

在近阈值计算(NTC)区域运行张量处理单元(TPU)可显著降低能耗,但会引入对工艺变化和数据活动的高度延迟敏感性。传统设计通常依赖保守的固定全局时钟来确保安全,这导致大部分时序裕量未被利用,因为大多数操作在时钟边沿之前很久就已完成。我们提出动态松弛感知时钟(DSAC),一种主动框架,用特定于操作的调整取代最坏情况时序。DSAC采用轻量级汉明距离、最高有效位和混合预测器来估计单个乘加(MAC)操作的延迟敏感性,并将其分为三个时序层级。这些层级通过固定全局参考时钟下的虚拟保持周期在本地强制执行,从而实现细粒度时序适应,无需全局时钟重调或频率缩放。闭环反馈控制器监控时序违规并在运行时更新层级阈值以保持鲁棒性。在量化DNN基准上的实验表明,MSB预测器保持了高推理精度,即使在激进性能点下平均损失仅为1%。此外,与基线TPU相比,DSAC在2.15倍频率缩放下实现了高达1.55倍的能效提升,同时面积开销低至13%。

英文摘要

Operating Tensor Processing Units (TPUs) in the near-threshold computing (NTC) region significantly reduces energy consumption but introduces high delay sensitivity to process variation and data activity. Conventional designs typically rely on a conservative, fixed global clock to ensure safety, which leaves large portions of timing margin unexploited as most operations finish well before the clock edge. We propose Dynamic Slack-Aware Clocking (DSAC), a proactive framework that replaces worst-case timing with operation-specific adjustments. DSAC employs lightweight Hamming-Distance, Most-Significant-Bit, and Hybrid predictors to estimate the delay sensitivity of individual multiply-accumulate (MAC) operations and classify them into three timing tiers. These tiers are enforced locally via dummy-hold cycles under a fixed global reference clock, enabling fine-grained timing adaptation without global clock retuning or frequency scaling. A closed-loop feedback controller monitors timing violations and updates tier thresholds at runtime to maintain resilience. Experiments on quantized DNN benchmarks demonstrate that the MSB predictor maintains high inference accuracy, with an average loss of only 1% even at aggressive performance points. Furthermore, DSAC achieves up to 1.55X better energy efficiency at 2.15X frequency scaling compared to a baseline TPU, while incurring an area overhead as low as 13%.

Comments31 pages, 13 figures. Accepted manuscript of ACM Transactions on Design Automation of Electronic Systems (TODAES); DOI: 10.1145/3837085

DOI:10.1145/3837085

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑