发表机构
University of Pisa; Istituto Nazionale di Ricerca Metrologica (INRiM); Politecnico di Torino(比萨大学; 国家计量研究院; 都灵理工大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出适配忆阻器的结构化无乘法器哈达玛储备池计算,以正交算子替代密集矩阵,在多硬件平台及27个基准测试中验证其性能,实现了高速低内存的规模化储备池计算。
AI 中文摘要
储备池计算(Reservoir Computing, RC)围绕固定(即未训练)的循环层设计循环神经网络,是神经形态硬件的天然候选方案。适配忆阻器的储备池从忆阻器器件动力学中衍生神经元动力学,但仍依赖密集循环矩阵,其物理实现成本高昂。本文用结构化正交算子替代密集矩阵,该算子由符号对角阵、置换矩阵和快速沃尔什-哈达玛变换构成,无乘法器,每步仅需O(N)个参数和O(NlogN)次运算,且从不以矩阵形式显式实现。我们将其实例化为标准及适配忆阻器的回声状态网络,每个单元含一个二元输入连接。数学分析表明,精确正交性在循环缩放中产生紧密的回声状态条件,且噪声响应可在设计时预测;此外,该算子单次应用即可混合整个状态。在20个分类和7个回归基准(储备池规模达N=8192)上的实验显示,结构化模型性能与密集正交储备池相当,且平均性能优于循环储备池,优势随规模增大而扩大。我们还在三种硬件平台上测试了循环步的耗时,其速度比密集乘积快达50倍,内存占用小10^4倍。最后,我们对该算子进行消融实验,并测量其对噪声、量化、器件失配及离散故障的响应。
英文摘要
Reservoir Computing (RC) designs Recurrent Neural Networks around a fixed, i.e., untrained, recurrent layer, and is a natural candidate for neuromorphic hardware. Memristive-friendly reservoirs derive the neuron dynamics from memristive-device kinetics, but still rely on dense recurrent matrices, which are expensive to realize physically. In this paper, we replace the dense matrix with a structured orthogonal operator, built from sign diagonals, a permutation, and a fast Walsh-Hadamard transform. The operator is multiplier-free, requires $O(N)$ parameters and $O(N\log N)$ operations per step, and is never materialized as a matrix. We instantiate it in a standard and in a memristive-friendly Echo State Network, with one binary input connection per unit. Our mathematical analysis shows that exact orthogonality yields an echo state condition that is tight in the recurrent scaling, and a noise response that is predictable at design time. Moreover, the operator mixes the whole state in a single application. Experiments on twenty classification and seven regression benchmarks, at reservoir sizes up to $N = 8192$, show that the structured models match dense orthogonal reservoirs, and achieve better mean performance than the cycle reservoir by a margin that widens with size. Furthermore, we time the recurrent step on three hardware platforms, where it is up to $50\times$ faster than a dense product and $10^4\times$ smaller in memory. Finally, we ablate the operator and measure the response to noise, quantization, device mismatch and discrete faults.
Commentssubmitted to Neurocomputing