EffVOC:无需相位的基于频谱表示的低延迟高效语音波形重建方法
EffVOC: Low-Delay Efficient Speech Waveform Reconstruction from Spectral Representations Without Phase
- Institute for Communications Technology, TU Braunschweig(通信技术研究所,布伦瑞克工业大学)
- Signal Processing Group, Universität Hamburg(信号处理组,汉堡大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
该研究提出EffVOC,一种低延迟高效声码器,可从幅度谱或梅尔系数输入合成宽带/全频带语音,在20 ms延迟下达到新SOTA,主观MOS分数接近真实值。
AI中文摘要:
Griffin-Lim算法是从幅度语谱图重建相位的开创性方法,但需要(无限)高算法延迟;其低延迟变体的语音质量较差。近期的(生成式)神经网络方法在中等至高算法延迟下提升了语音质量,但通常较为复杂且仅针对特定输入表示优化。我们基于一种高效低延迟声码器构建,提出EffVOC,其支持从幅度谱或梅尔系数输入合成宽带(WB)或全频带(FB)语音。我们在统一框架内针对多种模型规模评估这两种输入表示,并与当前最优方法(SOTA)对比。结果显示,我们提出的低延迟方法(20 ms,对比32 ms或更高)达到了新的SOTA,在幅度谱/梅尔表示下获得排名靠前的主观MOS分数(WB:4.17/4.15,FB:4.14/4.11),非常接近真实值。
英文摘要:
The Griffin-Lim algorithm has been a seminal contribution for phase reconstruction from amplitude spectrograms, however, requiring (infinitely) high algorithmic delay. Its low-delay variant suffers in speech quality. Recent (generative) neural network methods improve on speech quality still at medium to high algorithmic delay, but often they are complex and optimized only for one specific input representation. We build upon an efficient low-delay speech vocoder and propose EffVOC, which supports synthesis of wideband (WB) or fullband (FB) speech from either amplitude spectrum or Mel coefficient inputs. We evaluate both input representations across multiple model sizes in a unified framework and compare against state of the art. Results show that our proposed low-delay (20 ms vs. 32 ms or more) efficient approach marks a new SOTA by achieving top-ranked subjective MOS scores (WB: 4.17/4.15, FB: 4.14/4.11) for amplitude spectrum/Mel representations, very close to ground truth.