DualSpecSE:一种融合梅尔谱与复数谱的双路径语音增强网络
DualSpecSE: A Dual-Path Speech Enhancement Network Integrating Mel and Complex Spectrograms
浏览论文内容
中文总结 AI 辅助
提出DualSpecSE双路径语音增强网络,联合建模梅尔谱与复数谱,通过交互与融合模块提升ASR性能及语音重建质量。
中文摘要 AI 辅助
本文提出DualSpecSE,一种语音增强框架,采用双路径架构联合建模梅尔谱和复数谱,以提升ASR性能并实现更高质量的语音重建。梅尔分支学习粗粒度声学表征,生成增强梅尔谱以直接用于ASR;复数分支细化细粒度频谱细节,用于高保真波形重建。基于CleanMel的跨频带和窄频带模块,DualSpecSE引入交互模块和融合模块,以促进两个分支间的有效信息交换。该模型同时输出增强梅尔谱和复数谱,无需预训练声码器。实验结果表明,在语音保真度、感知质量和ASR性能方面均有一致提升。代码和音频样本已公开。
英文摘要
In this paper, we propose DualSpecSE, a speech enhancement framework that jointly models Mel-spectrogram and complex spectrogram in a dual-path architecture for improved ASR performance and higher-quality speech reconstruction. The Mel branch learns coarse-grained acoustic representations and produces enhanced Mel-spectrograms for direct ASR usage, while the complex branch refines fine-grained spectral details for high-fidelity waveform reconstruction. Built upon the cross-band and narrow-band blocks from CleanMel, DualSpecSE introduces an interaction module and a fusion module to enable effective information exchange between the two branches. The model simultaneously outputs enhanced Mel and complex spectrogram without requiring a pretrained vocoder. Experimental results demonstrate consistent improvements in speech fidelity, perceptual quality, and ASR performance. Codes and audio samples are available.
发表机构
- Northwestern Polytechnical University(西北工业大学)
- Huawei Technologies Co., Ltd.(华为技术有限公司)
机构由 AI 辅助整理,请以论文原文为准。