P2Flow:面向极端语音超分辨率的音素感知渐进流匹配
P2Flow: Phoneme-aware Progressive Flow Matching for Extreme Speech Super-Resolution
浏览论文内容
中文总结 AI 辅助
针对极端语音超分辨率中频谱输入严重受限的问题,提出音素感知渐进流匹配框架P2Flow,利用音素信息、渐进架构和声码器后训练,在TIMIT和VCTK上取得最先进结果。
中文摘要 AI 辅助
生成模型近期在语音超分辨率(SSR)领域展现出巨大潜力。然而,现有工作大多集中于标准或通用SSR配置,而严重受限频谱输入的极端设置在很大程度上仍未得到探索。在此场景下,现有方法表现出明显的性能退化,凸显了专门解决方案的必要性。为弥补这一空白,我们提出P2Flow,一个面向极端SSR的音素感知渐进流匹配(FM)框架,包含三项主要策略。首先,模型利用音素信息重建缺失的频谱成分。其次,采用渐进式架构设计,分层恢复不同频率区域。最后,引入声码器的后训练以增强整体波形保真度。我们在TIMIT和VCTK数据集上,于1 kHz至16 kHz及2 kHz至16 kHz两种设置下进行了大量实验,结果表明P2Flow在多项评估指标上取得了最先进的结果。
英文摘要
Generative models have recently demonstrated considerable promise in speech super-resolution (SSR). Nevertheless, the majority of existing work has concentrated on standard or versatile SSR configurations, leaving the extreme setting with severely limited spectral inputs largely unexplored. In this regime, current approaches exhibit marked performance degradation, underscoring the need for dedicated solutions. To bridge this gap, we introduce P2Flow, a phoneme-aware progressive flow matching (FM) framework designed for extreme SSR with three main strategies. First, our model leverages phonetic information to reconstruct missing spectral components. Furthermore, it employs a progressive architectural design that hierarchically restores distinct frequency regions. Finally, we incorporate post-training of the vocoder to enhance overall waveform fidelity. Extensive experiments are conducted on the TIMIT and VCTK datasets under both 1 kHz to 16 kHz and 2 kHz to 16 kHz settings, demonstrating that P2Flow yields state-of-the-art results across multiple evaluation metrics.
发表机构
- Stony Brook University(石溪大学)
- Northeastern University(东北大学)
- UIUC(伊利诺伊大学厄巴纳-香槟分校)
- UIC(伊利诺伊大学芝加哥分校)
机构由 AI 辅助整理,请以论文原文为准。