μNet:面向嵌入式数字信号处理器的超低内存、低复杂度语音增强模型
μNet: Ultra-Low-Memory and Low-Complexity Speech Enhancement for Embedded Digital Signal Processors
- International Audio Laboratories, Erlangen, Germany(国际音频实验室)
- Fraunhofer IIS, Erlangen, Germany(弗劳恩霍夫集成电路研究所)
- Friedrich-Alexander-Universität Erlangen-Nürnberg (FAU), Germany(埃尔朗根-纽伦堡大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
针对嵌入式DSP语音增强的多约束需求,提出μNet端到端DNN模型,其内存、复杂度、延迟均极低,可兼容神经加速器并支持消费级DSP的全整数运算,性能与最优方法相当。
AI中文摘要:
嵌入式数字信号处理器(DSP)上的语音增强对内存占用、计算复杂度、延迟以及整数运算支持有严格要求。尽管近期基于深度神经网络(DNN)的方法已分别应对了这些挑战,但现有文献中尚无统一框架能同时满足所有要求以实现实际部署。本研究提出μNet,一款超低内存、低复杂度、低延迟的端到端DNN模型。该方法仅需90KB静态内存与28MMACs,同时支持低至4ms的算法延迟,性能与复杂度相近的现有最优方法相当。实验表明,μNet可兼容神经加速器,并在Cadence Tensilica HiFi 4/5等消费级DSP平台上支持全整数算术运算。
英文摘要:
Speech enhancement on embedded digital signal processors (DSPs) imposes strict constraints on memory footprint, computational complexity, latency, and support for integer operations. Although recent DNN-based approaches have addressed these challenges individually, no unified framework in the literature simultaneously addresses all these requirements for practical deployment. In this work, we propose μNet, an ultra-low-memory, low-complexity, and low-latency end-to-end DNN model. The proposed method requires only $90$~KB of static memory and $28$~MMACs, while supporting an algorithmic latency as low as $4$~ms with performance comparable to state-of-the-art methods of similar complexity. Our experiments demonstrate that μNet is compatible with neural accelerators and supports full integer-arithmetic operations on consumer DSP platforms such as Cadence Tensilica HiFi 4/5.