AI 中文总结
本文提出DuSpaR,一种双状态稀疏循环单元,通过ReLU稀疏化和反馈调制降低计算量,在KWS、SLU和SE任务上以更少计算实现更高或相当的性能。
AI 中文摘要
我们引入了双状态稀疏循环单元(DuSpaR),作为资源受限边缘设备上语音处理模型的一种计算高效的构建模块。它采用双状态循环在状态反馈回路中调制其输入向量。其循环单元使用ReLU激活函数对矩阵-向量乘法中涉及的输入向量操作数进行稀疏化。通过动态跳过零条目,可以在推理时节省乘加运算和权重内存读取。我们在三个语音任务上评估了DuSpaR:在Google Speech Commands数据集上的关键词识别(KWS)、在Fluent Speech Commands数据集上的口语语言理解(SLU),以及在Voice Bank + Demand(VBD)数据集上的语音增强(SE)。在相似的参数数量下,DuSpaR在KWS和SLU上分别比门控循环单元(GRU)少51.0%和68.4%的计算量,同时实现更高的准确率;在SE上少50.1%的计算量,同时保持相似的质量。在可比较的计算成本下,跨一系列模型规模,DuSpaR也实现了比其它稀疏感知循环模型更高的KWS/SLU准确率和更好的SE质量。消融研究表明,与单状态循环基线相比,双状态循环在相似任务性能下将有效计算减少了3.1到11.3倍。
英文摘要
We introduce the Dual-state Sparsifying Recurrent Unit (DuSpaR) as a computationally efficient building block for speech processing models on resource-constrained edge devices. It employs dual-state recurrence to modulate its input vectors in a stateful feedback loop. Its recurrent cells sparsify the input vector operand involved in matrix-vector multiplication using ReLU activation. By skipping the zero entries dynamically, inference-time savings in multiply-accumulate operations and weight memory fetches can be achieved. We evaluate DuSpaR on three speech tasks: keyword spotting (KWS) on the Google Speech Commands dataset, spoken language understanding (SLU) on the Fluent Speech Commands dataset, and speech enhancement (SE) on the Voice Bank + Demand (VBD) dataset. At similar parameter counts, DuSpaR requires 51.0% and 68.4% less computation than Gated Recurrent Unit (GRU) on KWS and SLU, respectively, while achieving higher accuracy, and 50.1% less computation on SE while maintaining similar quality. At comparable computational cost and across a range of model sizes, DuSpaR also achieves higher KWS/SLU accuracy and better SE quality than other sparsity-aware recurrent models. Ablation studies show that compared with the single-state recurrence baseline, dual-state recurrence reduces the effective compute by factors of 3.1 to 11.3 at similar task performance.