AI 中文总结
研究张量加速器华莱士树乘法器随机活动预测问题,提出随机活动预测(SAP)方法,通过检查操作数汉明权重等进行预测,有谱收缩引理等三个形式化结果支持,解决了现有技术未覆盖的功率堆栈特定层问题。
AI 中文摘要
张量加速器乘法器即使在稀疏操作数几乎不需要内部切换时,每个时钟周期也会消耗动态功率。现有技术无法解决此问题:零检测要求操作数完全为零,结构功率门控要求乘法器空闲,离线权重选择无法响应运行时数据。本文介绍了随机活动预测(SAP),通过在乘法器执行前检查输入操作数的汉明权重,预测低切换活动,并在确定性安全控制器独立确认复用正确时冻结输入来解决此问题。误预测只会导致节省缺失,不会产生错误答案。有三个形式化结果支持SAP:(i)一个谱收缩引理,证明华莱士树活动取决于操作数位密度而非位位置,在256周期窗口内建立了利普希茨常数$L\phi = 3/2$且预测误差低于$10^{-13}$;(ii)一个信息保留定理,表明$\eta_I \ge 1 - O(\log n/n)$,所以每周期一位捕获了关于$O(n^2)$个内部节点几乎所有的预测信息;(iii)一个伯努利最优性定理,证明在所考虑的汉明权重统计的校准一位编码器家族中,所选编码是最优的。SAP解决了现有技术未涵盖的张量加速器功率堆栈的特定层问题。
英文摘要
Tensor accelerator multipliers burn dynamic power on every clock cycle, even when sparse operands require very little internal switching. No existing technique addresses this: zero-detection requires exactly-zero operands, structural power gating requires an idle multiplier, and offline weight selection cannot respond to runtime data. This paper introduces Stochastic Activity Prediction (SAP), which closes this gap by examining the Hamming weight of arriving operands before the multiplier executes, predicting low switching activity, and freezing the inputs when a deterministic Safety Controller independently confirms the reuse is correct. Mispredictions cause missed savings, never wrong answers. Three formal results underpin SAP: (i) a Spectral Contraction Lemma proving that Wallace-tree activity depends on operand bit density, not bit position, establishing Lipschitz constant $Lϕ= 3/2$ and prediction error below $10^{-13}$ for a 256-cycle window; (ii) an Information Retention Theorem showing $η_I \ge 1 - O(\log n/n)$, so one bit per cycle captures nearly all predictive information about $O(n^2)$ internal nodes; and (iii) a Bernoulli Optimality Theorem proving the chosen encoding is shown to be optimal, within the family of calibrated one-bit encoders of Hamming-weight statistics considered. SAP addresses the specific layer of the tensor accelerator power stack that existing techniques do not cover.