arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.12082eess.AS

在神经音频编解码器的隐空间中重新思考基于语言模型的生成式语音增强

Rethinking Language Model-Based Generative Speech Enhancement in the Latent Space of a Neural Audio Codec

Yihui Fu, Zhengyang Li, Tim Fingscheidt

首次发表
浏览论文内容

中文总结 AI 辅助

本文提出涵盖六种基于LM的生成式SE范式的统一框架,对比其性能并提出辅助损失微调策略,发现连续域范式优于离散域,CNAR效果最佳且微调策略可提升多指标。

中文摘要 AI 辅助

基于语言模型(LM)的语音增强(SE)近年来借助神经音频编解码器(NAC)的隐空间特征快速发展。本文首先提出一个统一框架,涵盖基于离散/连续隐式NAC特征的六种流行基于LM的生成式SE建模范式:离散或连续自回归(D/CAR)SE、离散或连续非自回归(D/CNAR)SE、离散扩散(DDiff)SE以及连续流匹配(CFM)SE。其次,我们首次在统一实验设置与概述中,采用多种侵入式和非侵入式指标对比它们的性能,以实现公平且全面的评估。第三,我们提出一种带辅助损失的微调策略,用于重构语音,以同时提升侵入式和非侵入式指标。在URGENT 2025语音增强挑战赛数据划分上进行训练与评估后,所有连续域范式均优于离散域对应范式,整体最佳方法为CNAR。我们进一步表明,所提出的辅助损失微调策略可在全部六种范式中持续提升DNSMOS、NISQA、PESQ和POLQA指标。

英文摘要

Language model (LM)-based speech enhancement (SE) has recently emerged rapidly using latent space features of neural audio codecs (NACs). In this paper, first, we present a unified framework covering six popular LM-based generative SE modeling paradigms based on discrete/continuous latent NAC features: discrete or continuous autoregressive (D/CAR) SE, discrete or continuous non-autoregressive (D/CNAR) SE, discrete diffusion (DDiff) SE, and continuous flow matching (CFM) SE. Second, we are the first to compare their performance in a unified experimental setup and synopsis with diverse intrusive and non-intrusive metrics, enabling a fair and comprehensive evaluation. Third, we propose a fine-tuning strategy with auxiliary losses on reconstructed speech to improve both intrusive and non-intrusive metrics. Trained and evaluated on URGENT 2025 Speech Enhancement Challenge data splits, all continuous-domain paradigms excel their discrete-domain counterparts. The overall best approach turns out to be CNAR. We further show that our proposed auxiliary loss fine-tuning strategy helps to improve DNSMOS, NISQA, PESQ, and POLQA consistently in all six paradigms.

补充信息

↑