基于Transformer的神经波束成形,用于智能低功耗可听戴设备的实时语音增强
Transformer-based Neural Beamforming for Real-Time Speech Enhancement on Smart Low-Power Hearable Devices
- University of Bologna(博洛尼亚大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
提出一种在低功耗MCU上实时运行基于Transformer的MVDR神经波束成形方法,通过混合精度与时间切片实现15毫秒延迟,达到高语音质量并显著延长电池寿命。
AI中文摘要:
准确、高效且低延迟的空间波束成形是新兴智能可听戴设备中的关键组件,能够在抑制噪声和干扰的同时增强语音。然而,在严格的实时约束下处理多个输入源,对可听戴设备中使用的低功耗、资源受限的微控制器单元(MCU)构成了重大挑战。我们提出了一种在MCU上实时执行基于神经网络的最小方差无失真响应(MVDR)波束成形器的优化方法。利用六个麦克风和三级混合精度方案(float32 MVDR、int8 CNN、float16 Transformer),该流水线将估计MVDR权重的CNN与执行逐帧校正的轻量级Transformer配对。通过将权重估计与波束成形进行时间切片,实现了15毫秒的逐帧延迟,同时每564毫秒刷新一组完整的CNN推导权重。部署的混合精度流水线在平均功耗45.9毫瓦下,达到了97.65%的短时客观可懂度(STOI)、20.26分贝的尺度不变信噪比(SI-SNR)和3.676的宽带PESQ。语音活动检测(SAD)模块(准确率98.5%,每次推理0.62毫焦耳)在静音期间绕过流水线;在现实部署条件下,该系统在100毫安时电池上超过了16小时的全天目标,估计寿命可达约20小时。据我们所知,这是首个在MCU级设备上部署的实时多通道基于Transformer的神经波束成形流水线。
英文摘要:
Accurate, efficient, and low-latency spatial beamforming is a key component in emerging smart hearable devices, enhancing speech while suppressing noise and interference. However, handling multiple input sources under strict real-time constraints poses significant challenges for the low-power, resource-constrained microcontroller units (MCUs) used in hearables. We present an optimized methodology for the real-time execution of a neural-network-based minimum variance distortionless response (MVDR) beamformer on MCUs. Using six microphones and a three-stage mixed-precision scheme (float32 MVDR, int8 CNN, float16 Transformer), the pipeline pairs a CNN that estimates the MVDR weights with a lightweight Transformer that applies a per-frame correction. By time-slicing weight estimation with beamforming, it achieves a 15~ms per-frame latency while refreshing a complete set of CNN-derived weights every 564~ms. The deployed mixed-precision pipeline attains a short-time objective intelligibility (STOI) of 97.65\%, a scale-invariant signal-to-noise ratio (SI-SNR) of 20.26~dB, and a wideband PESQ of 3.676 at an average power of 45.9~mW. A speech activity detection (SAD) module (98.5\% accuracy, 0.62~mJ per inference) bypasses the pipeline during silence; under realistic deployment conditions, the system exceeds the 16~h all-day target on a 100~mAh battery, with an estimated lifetime of up to $\sim$20~h. To our knowledge, this is the first real-time multi-channel Transformer-based neural beamforming pipeline deployed on an MCU-class device.