arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

StreamWSR:可流式且轻量的波形域神经语音超分辨率模型

StreamWSR: Streamable and Lightweight Waveform-Domain Neural Speech Super-Resolution

Yuan Tian, Yang Ai, Hui-Peng Du, Zhen-Hua Ling

arXiv 2609.03381首次发表:更新:

AI 中文总结

本文提出StreamWSR这一可流式轻量波形域语音超分辨率模型,通过因果架构实现零前瞻流式推理,在16 kHz语音超分辨率任务上,以9M参数、2G FLOPs达到优于或相当的语音质量与可懂度。

AI 中文摘要

本文提出了StreamWSR,一种用于语音超分辨率(SR)的可流式神经波形域模型。通过采用具有紧凑帧级波形表示的全因果架构,所提出的StreamWSR支持零前瞻流式推理,同时避免基于声码器的重建和显式相位预测。具体而言,StreamWSR通过步长因果卷积将输入波形下采样为紧凑的帧级表示;随后,采用轻量的因果长短时建模主干,在因果约束下捕获局部波形结构和长程历史依赖;最后,通过因果转置卷积将建模输出转换回波形域,并通过残差连接与输入波形结合,生成最终的高分辨率语音。在16 kHz语音SR上的实验结果表明,与代表性的波形域和频谱域基线相比,StreamWSR实现了具有竞争力或更优的语音质量和可懂度,同时保持零前瞻流式优势,仅需9M参数和2G FLOPs。

英文摘要

This paper proposes StreamWSR, a Streamable neural Waveform-domain model for speech Super-Resolution (SR). By adopting a fully causal architecture with compact frame-level waveform representation, the proposed StreamWSR supports zero-look-ahead streaming inference while avoiding vocoder-based reconstruction and explicit phase prediction. Specifically, StreamWSR downsamples the input waveform into a compact frame-level representation using strided causal convolutions. Then, a lightweight causal long-short-term modeling backbone is employed to capture both local waveform structures and long-range historical dependencies under causal constraints. Finally, the modeled output is converted back to the waveform domain through a causal transposed-convolution and combined with the input waveform via a residual connection to generate the final high-resolution speech. Experimental results on 16 kHz speech SR show that StreamWSR achieves competitive or superior speech quality and intelligibility compared with representative waveform- and spectrum-based baselines, while maintaining a zero-look-ahead streaming advantage with only 9M parameters and 2G FLOPs.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑