arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

PhaseGAN:通过解耦幅度和GAN驱动的相位重建实现高保真声码器

PhaseGAN: High-Fidelity Vocoder via Decoupled Amplitude and GAN-Driven Phase Reconstruction

Wenzheng Zhang, Xueliang Zhang, Shulin He, Fei Zhao, Xin Liu, Pengjie Shen, Zhenlong Guo, Zixuan Xue, Hongtao Bao, Zixuan Li

arXiv 2609.12918首次发表:更新:

发表机构

Inner Mongolia University(内蒙古大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

PhaseGAN通过解耦幅度与相位重建的轻量级声码器,以少量参数实现高保真音频合成,并展现跨领域泛化能力。

AI 中文摘要

声码器是现代文本到语音(TTS)系统的关键组成部分。尽管基于神经网络的声码器取得了显著进展,但精确的相位重建仍然是限制音频质量和建模效率的主要挑战。我们引入了PhaseGAN,一种轻量级声码器,通过“梅尔频谱→幅度→相位”的重建流程来解决这一限制。通过不同的方法重建幅度和相位谱,所提出的PhaseGAN在利用更少的模型参数和降低的计算需求的同时,超越了最先进的基线。其紧凑版本以约500K参数和1 GMAC的计算负载生成高保真音频,使其非常适合边缘设备上的实时应用。此外,我们的方法尽管没有在音乐数据上进行训练,却展现出卓越的音乐音频合成能力,体现了前所未有的跨领域泛化能力。请参见此URL以获取我们工作的演示。

英文摘要

A vocoder is a pivotal component of modern text-to-speech (TTS) systems. Despite the significant progress of neural network-based vocoders, accurate phase reconstruction remains the main challenge limiting both audio quality and modeling efficiency. We introduce PhaseGAN, a lightweight vocoder that addresses this limitation through a "mel $\rightarrow$ Amplitude $\rightarrow$ Phase" reconstruction pipeline. By reconstructing amplitude and phase spectra via distinct methodologies, the proposed PhaseGAN outperforms state-of-the-art baselines while utilizing fewer model parameters and reduced computational requirements. The compact version generates high-fidelity audio with approximately 500K parameters and 1 GMAC computational load, making it highly suitable for real-time applications on edge devices. In addition, our approach exhibits exceptional musical audio synthesis capabilities despite no training on musical data, illustrating unprecedented cross-domain generalization. See https://github.com/phasegan/phasegan-audio-demo for demos of our work.

Comments15 pages, 1 figure

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑