arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.29866eess.AScs.LGcs.SD

超越模型规模:重新设计LiSenNet用于嵌入式语音增强

Beyond Model Size: Redesigning LiSenNet for embedded speech enhancement

  • GN A/S(GN集团)

机构由 AI 辅助整理,请以论文原文为准。

Clément Laroche, Rasmus Kongsgaard Olsson

AI总结:

针对受限NPU上实时语音增强的算子不兼容问题,重新设计37k参数LiSenNet为静态int8兼容模型,在STM32N6570上达到PESQ 3.08(FP32)和实时因子0.30,证明参数、算子、量化与流状态需协同设计。

AI中文摘要:

在资源受限设备上部署实时语音增强需要满足严格的延迟、内存和能量约束。微控制器NPU可以在这些约束下加速神经推理,但只能通过静态、整数量化图中的受限算子集合来实现。最近的语音增强网络已将参数数量和MACs降低到名义上适合微控制器的水平,但其算子和执行模式通常仍与受限NPU不兼容。我们通过为STM32N6570-DK Neural-ART加速器重新设计LiSenNet(一个37k参数的子带双路径模型)来解决这一差距。我们用卷积频率和时间混合器替换其循环瓶颈,将不支持的操作重新表述为静态int8兼容原语,并使用有界解码器激活来保持量化后的质量。在VoiceBank-DEMAND上,最终的NPU兼容模型达到或超过循环LiSenNet基线,在FP32中达到PESQ 3.08对比3.01,在int8中达到3.01对比2.93。部署在微控制器上,它处理每个16毫秒输入跳跃耗时4.83毫秒,对应实时因子为0.30。无状态感受野重计算在相同帧率下慢一个数量级,尽管加速器利用率更高。这些结果表明,参数数量和算子兼容性、量化范围以及持久流状态必须共同设计,才能在受限NPU上实现高效的实时语音增强。

英文摘要:

Deploying real-time speech enhancement on resource-constrained devices requires meeting strict latency, memory, and energy constraints. Microcontroller NPUs can accelerate neural inference under these constraints, but only through a restricted set of operators in static, integer-quantized graphs. Recent speech-enhancement networks have reduced parameter counts and MACs to levels nominally suitable for microcontrollers, but their operators and execution patterns often remain incompatible with restricted NPUs. We address this gap by redesigning LiSenNet, a 37k parameter sub-band dual-path model, for the STM32N6570-DK Neural-ART accelerator. We replace its recurrent bottleneck with convolutional frequency and temporal mixers, reformulate unsupported operations as static int8-compatible primitives, and use bounded decoder activations to preserve quality after quantization. On VoiceBank-DEMAND, the final NPU-compatible model matches or exceeds the recurrent LiSenNet baseline, reaching PESQ 3.08 versus 3.01 in FP32 and 3.01 versus 2.93 in int8. Deployed on a microcontroller, it processes each 16 ms input hop in 4.83 ms, corresponding to a real-time factor of 0.30. Stateless receptive-field recomputation is an order of magnitude slower at the same frame rate despite higher accelerator utilization. These results show that parameter count and operator compatibility, quantization range, and persistent streaming state must be co-designed to achieve efficient real-time speech enhancement on restricted NPUs.

↑