arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

基于切比雪夫多项式的非线性激活函数与Softmax计算的脉动阵列架构

A Systolic Array Architecture for Nonlinear Activation Functions and Softmax Computation using Chebyshev Polynomials

Benedikt Schaible, Anirudh Suresh Bharadwaj, Ulf Schlichtmann, Jiang Hu

arXiv 2608.04734首次发表:更新:

AI 中文总结

本文提出一种基于切比雪夫多项式的脉动阵列激活单元架构,可支持多种单变量激活函数与Softmax函数,相比CORDIC等方案,在误差、面积和功耗上均有显著优化,实现了硬件资源的共享。

AI 中文摘要

近年来,神经网络加速器因比基于CPU的平台效率更高而受到广泛关注。这些加速器通常使用不同的硬件单元处理单变量激活函数(如tanh)和多变量Softmax函数,从而错失了两者间资源共享的机会。本文提出了一种新型的基于脉动阵列的激活单元架构,该架构支持多种单变量激活函数及Softmax函数。通过利用切比雪夫多项式近似,与CORDIC基线相比,我们的激活函数单元在tanh上的平均绝对误差降低了71%,同时面积减少4.6%,功耗降低5.1%;与CORDIC和分段线性近似相比,我们的Softmax近似的KL散度分别降低了44.6%和79.0%。

英文摘要

Neural Network Accelerators have gained popularity in recent years due to their greater efficiency than CPU-based platforms. Often, these accelerators utilize different hardware units for univariate activation functions, such as tanh, and the multivariate softmax, thereby missing opportunities for resource sharing between them. In this paper, we describe a novel systolic array-based activation unit architecture that supports multiple univariate activation functions as well as the softmax function. By utilizing Chebyshev polynomial approximations, our activation function unit achieves up to 71% lower mean absolute error for tanh compared to a CORDIC baseline, while using 4.6% less area and 5.1% less power. Our softmax approximation enables a 44.6% and 79.0% lower KL divergence compared to CORDIC and a piecewise-linear approximation, respectively.

CommentsAccepted for IEEE COINS 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑