arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.13911eess.AScs.SD

DualSpecSE:一种融合梅尔谱与复数谱的双路径语音增强网络

DualSpecSE: A Dual-Path Speech Enhancement Network Integrating Mel and Complex Spectrograms

Xingchen Li, Ziqian Wang, Zikai Liu, Yike Zhu, Zihan Zhang, Longshuai Xiao, Lei Xie

首次发表
浏览论文内容

中文总结 AI 辅助

提出DualSpecSE双路径语音增强网络,联合建模梅尔谱与复数谱,通过交互与融合模块提升ASR性能及语音重建质量。

中文摘要 AI 辅助

本文提出DualSpecSE,一种语音增强框架,采用双路径架构联合建模梅尔谱和复数谱,以提升ASR性能并实现更高质量的语音重建。梅尔分支学习粗粒度声学表征,生成增强梅尔谱以直接用于ASR;复数分支细化细粒度频谱细节,用于高保真波形重建。基于CleanMel的跨频带和窄频带模块,DualSpecSE引入交互模块和融合模块,以促进两个分支间的有效信息交换。该模型同时输出增强梅尔谱和复数谱,无需预训练声码器。实验结果表明,在语音保真度、感知质量和ASR性能方面均有一致提升。代码和音频样本已公开。

英文摘要

In this paper, we propose DualSpecSE, a speech enhancement framework that jointly models Mel-spectrogram and complex spectrogram in a dual-path architecture for improved ASR performance and higher-quality speech reconstruction. The Mel branch learns coarse-grained acoustic representations and produces enhanced Mel-spectrograms for direct ASR usage, while the complex branch refines fine-grained spectral details for high-fidelity waveform reconstruction. Built upon the cross-band and narrow-band blocks from CleanMel, DualSpecSE introduces an interaction module and a fusion module to enable effective information exchange between the two branches. The model simultaneously outputs enhanced Mel and complex spectrogram without requiring a pretrained vocoder. Experimental results demonstrate consistent improvements in speech fidelity, perceptual quality, and ASR performance. Codes and audio samples are available.

发表机构

  • Northwestern Polytechnical University(西北工业大学)
  • Huawei Technologies Co., Ltd.(华为技术有限公司)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑