AI 中文总结
本文提出PENDA处理单元,利用余弦定理将乘法转为平方差运算,在保持精确性的同时,使深度学习加速器PE阵列的面积、能耗和时钟周期分别减少11%~36%、5%~48%和11%~19%。
AI 中文摘要
内积计算主导了深度学习模型的计算成本;因此,加速这一基本操作是提高硬件效率的关键。然而,大多数现有技术依赖于近似方法,这可能会降低模型精度。为了在优化硬件的同时保持精确性,本文提出了PENDA(基于差范数架构的处理单元),它利用余弦定理将乘法重新表述为平方差运算。用所提出的差范数单元替换乘累加单元,在深度学习加速器的处理单元阵列中,面积、能耗和时钟周期分别减少了11%至36%、5%至48%和11%至19%。
英文摘要
Inner product computation dominates the computational cost of deep learning models; thus, accelerating this primitive is key to improving hardware efficiency. However, most existing techniques rely on approximations, which can degrade model accuracy. To preserve exactness while optimizing hardware, this paper presents PENDA (processing element via norm-of-difference architecture), which leverages the law of cosines to recast multiplications as squared-difference operations. Replacing multiply-accumulate units with the proposed norm-of-difference units yields 11~36%, 5~48%, and 11~19% reductions in area, energy, and clock period, respectively, for the PE array of a deep learning accelerator.
Commentssubmitted to IEEE TCAS-AI