逐帧提前退出是否划算?面向设备端语音增强的动态深度计算匹配研究
Does per-frame early exit pay? A compute-matched study of dynamic depth for on-device speech enhancement
- GN A/S(GN公司)
- Technical University of Denmark(丹麦技术大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本研究提出一种监督训练协议,使因果模型各深度输出不劣于浅层,从而生成更高效的静态模型,在设备端语音增强中实现等效计算下PESQ提升0.11,并以30%更少计算匹配最佳PESQ,动态执行开销微小。
AI中文摘要:
基于深度学习的语音增强技术正越来越多地部署在助听器、头戴式耳机和耳塞等设备端。然而,这些设备中的大多数只能加速静态int8图,因此深度可变的网络必须实现为多个图,并由一个策略进行编排。在本文中,我们对一个因果模型的每个中间深度进行监督训练,然后微调其输出头,以确保更深层的输出永远不会比更浅层的输出差。利用这一训练协议,我们可以得到一系列静态模型,这些模型在同等预算下比从零开始训练的同等规模模型具有更高的帕累托效率。具体而言,在等效计算量下,我们实现了高达0.11的PESQ提升,并以30%更少的计算量匹配了最佳PESQ。随后,我们将模型量化到int8,并在STM32N6微控制器上测量延迟-质量前沿。在VoiceBank-DEMAND数据集上,动态增强器与静态模型处于同一前沿,而不是以质量换取动态执行。在配套的Cortex-M55上运行该策略每帧仅需26微秒,而将增强器拆分为独立的NPU图则增加了2.2%的延迟开销。因此,动态执行的代价很小。
英文摘要:
Deep learning-based speech enhancement is increasingly deployed on-device in hearing aids, headsets, and earbuds. Most of these devices, however, can only accelerate static int8 graphs, so a depth-varying network must be implemented as several graphs, orchestrated by a policy. In this paper, we supervise every intermediate depth of one causal model, then we fine-tune its output heads to guarantee that deeper outputs are never worse than shallower ones. Using this training protocol, we can derive a family of static models that are more Pareto-efficient than their equivalently-sized counterparts trained from scratch on the same budget. Specifically, we achieve up to 0.11 higher PESQ for equivalent compute, and match the best PESQ at 30% less compute. We then quantize the models to int8 and measure the latency-quality frontier on an STM32N6 microcontroller. On VoiceBank-DEMAND, the dynamic enhancer lies on the same frontier as the static models, rather than trading quality for dynamic execution. Running the policy on the companion Cortex-M55 takes only 26 $μ$s per frame, while splitting the enhancer into separate NPU graphs adds 2.2% latency overhead. The cost of dynamic execution is therefore small.