SCNet:利用子带条件网络和幅度感知相位损失增强基于GAN的语音生成
SCNet: Enhancing GAN-based Speech Generation with Subband Condition Network and Magnitude-aware Phase Loss
浏览论文内容
中文总结 AI 辅助
SCNet提出子带条件网络与幅度感知相位损失,增强GAN声码器,提升语音生成质量。
中文摘要 AI 辅助
最近的语音生成主要依赖于基于GAN的网络,旨在从梅尔频谱图合成高质量波形。然而,这些方法通常作为黑盒模型运行,导致固有频谱信息的丢失。在这项工作中,我们提出了SCNet,一种增强子带条件网络的基于GAN的声码器,以解决这一问题。具体来说,SCNet利用轻量级条件网络预测的子带信号作为先验知识。然后通过STFT变换该子带信号以获得傅里叶系数,并将其集成到主干网络中以增强重建。此外,为了缓解相位缠绕,我们引入了幅度感知相位损失,该损失计算由相应幅度加权的瞬时相位误差,强调能量较高的区域。实验结果表明,SCNet在高品质语音生成的客观和主观评估中均实现了优越的性能。
英文摘要
Recent speech generation has been predominantly driven by GAN-based networks aimed at high-quality waveform synthesis from mel-spectrograms. However, these methods often operate as black-box models, leading to the loss of inherent spectral information. In this work, we propose SCNet, a GAN-based vocoder augmented with a Subband Condition Network to address this issue. Specifically, SCNet leverages a subband signal predicted by a lightweight condition network as prior knowledge. This subband signal is then transformed via STFT to obtain Fourier coefficients, which are integrated into the backbone for the enhanced reconstruction. Additionally, to mitigate the phase wrapping, we introduce a magnitude-aware phase loss that computes instantaneous phase errors weighted by the corresponding magnitude, emphasizing regions with higher energy. Experimental results demonstrate that SCNet achieves superior performance in both objective and subjective evaluations for high-quality speech generation.