用于脉冲密度调制语音信号端到端处理的多速率状态空间模型
Multirate State Space Models for End-to-End Processing of Pulse Density Modulated Speech Signals
浏览论文内容
中文总结 AI 辅助
该研究提出多速率状态空间模型,构建PDM语音端到端处理架构,可实现调制与采样率无关的表示,在不同采样率下性能优异,大幅降低处理时间步数。
中文摘要 AI 辅助
基于状态空间模型(SSM)的深度神经网络(DNN)正越来越多地应用于语音处理,但通常处理的是脉冲编码调制(PCM)音频。这限制了其在低功耗、常在线边缘设备上的部署,这类设备通常使用单比特脉冲密度调制(PDM)微机电(MEMS)麦克风,因其具有抗噪性强、成本低、采样率可变以支持低功耗运行的优势。实际上,将PDM转换为PCM需要低通滤波和降采样,这对资源受限的硬件会带来高昂开销。尽管已有研究尝试直接处理PDM信号,但这些方法需要较长的训练时间,且在不同采样率下泛化性较差。本文表明,SSM具备两个可解决这些问题的关键特性:其连续时间参数化使其能够生成输入音频信号的一致表示,而不受调制策略和采样率的影响;其长期记忆能力使该表示可被大幅降采样,无需任何抗混叠操作。随后,我们提出了一种新颖的PDM语音端到端处理架构,该架构使用SSM将输入音频信号编码为与调制方式和采样率无关的潜在表示。我们的实验表明,所提架构在低功耗采样率(512 kHz)下实现了稳健的语音分类与增强性能,在标准PDM采样率(2 MHz)下测试时,其性能与处理PCM数据的最先进算法相当。此外,我们还表明,SSM的输出可被降采样超过65000倍,从而显著减少下游层的处理时间步数。
英文摘要
Deep neural networks (DNNs) based on state-space models (SSMs) are increasingly applied to speech processing, but typically operate on pulse-code-modulated (PCM) audio. This constrains deployment on low-power, always-on edge devices, which commonly use single-bit pulse-density-modulated (PDM) micro-electromechanical (MEMS) microphones for their noise robustness, low cost, and variable sampling rates that enable low-power operation. In fact, converting PDM to PCM requires low-pass filtering and decimation, imposing costly overhead on resource-constrained hardware. While prior works have attempted to process PDM signals directly, they require long training times and generalize poorly across sampling rates. In this paper, we show that the SSM has two key properties that remediate these issues: its continuous-time parametrization allows it to produce a consistent representation of the input audio signal, regardless of the modulation strategy and sampling rate, and its long-term memory enables this representation to be aggressively downsampled without needing any anti-aliasing operations. We then propose a novel end-to-end PDM speech processing architecture that uses an SSM to encode the input audio signal into a modulation- and sampling-rate-invariant latent representation. We show that our proposed architecture achieves robust speech classification and enhancement gains at low-power sampling-rates (512 kHz) and similar performance to state-of-the-art algorithms operating on PCM data when tested on standard PDM sampling-rates of 2 MHz. Moreover, we show that the SSM's output can be downsampled by more than 65,000 times, thus significantly reducing the number of processing timesteps in downstream layers.
发表机构
- University of Sherbrooke(谢布鲁克大学)
机构由 AI 辅助整理,请以论文原文为准。