发表机构
Inner Mongolia University(内蒙古大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
PhaseGAN通过解耦幅度与相位重建的轻量级声码器,以少量参数实现高保真音频合成,并展现跨领域泛化能力。
AI 中文摘要
声码器是现代文本到语音(TTS)系统的关键组成部分。尽管基于神经网络的声码器取得了显著进展,但精确的相位重建仍然是限制音频质量和建模效率的主要挑战。我们引入了PhaseGAN,一种轻量级声码器,通过“梅尔频谱→幅度→相位”的重建流程来解决这一限制。通过不同的方法重建幅度和相位谱,所提出的PhaseGAN在利用更少的模型参数和降低的计算需求的同时,超越了最先进的基线。其紧凑版本以约500K参数和1 GMAC的计算负载生成高保真音频,使其非常适合边缘设备上的实时应用。此外,我们的方法尽管没有在音乐数据上进行训练,却展现出卓越的音乐音频合成能力,体现了前所未有的跨领域泛化能力。请参见此URL以获取我们工作的演示。
英文摘要
A vocoder is a pivotal component of modern text-to-speech (TTS) systems. Despite the significant progress of neural network-based vocoders, accurate phase reconstruction remains the main challenge limiting both audio quality and modeling efficiency. We introduce PhaseGAN, a lightweight vocoder that addresses this limitation through a "mel $\rightarrow$ Amplitude $\rightarrow$ Phase" reconstruction pipeline. By reconstructing amplitude and phase spectra via distinct methodologies, the proposed PhaseGAN outperforms state-of-the-art baselines while utilizing fewer model parameters and reduced computational requirements. The compact version generates high-fidelity audio with approximately 500K parameters and 1 GMAC computational load, making it highly suitable for real-time applications on edge devices. In addition, our approach exhibits exceptional musical audio synthesis capabilities despite no training on musical data, illustrating unprecedented cross-domain generalization. See https://github.com/phasegan/phasegan-audio-demo for demos of our work.
Comments15 pages, 1 figure