arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.03940cs.SDcs.AI

基于连续神经音频编解码器表示的掩码自回归语音增强

Masked Autoregressive Speech Enhancement with Continuous Neural Audio Codec Representations

Yoto Fujita, Simon Leglaive, Laurent Girin

首次发表
浏览论文内容

中文总结 AI 辅助

该研究提出MARSE方法,利用连续NAC表示迭代解码掩码纯净语音帧,可灵活权衡语音增强性能与计算成本。

中文摘要 AI 辅助

以往大多数基于掩码生成建模的语音增强(SE)工作依赖于使用神经音频编解码器(NAC)获得的音频信号离散令牌表示。然而,近期一项研究表明,NAC的连续潜在表示在语音质量和可懂度方面对SE更具优势。本研究提出掩码自回归语音增强(MARSE)方法,该方法基于语音的连续NAC表示迭代解码掩码纯净语音帧。具体而言,在控制其他条件相同的情况下,即使用相同的DNN(Conformer模型)、相同的NAC(DAC编解码器)和相同的训练设置,研究人员考察了一组不同的解码策略。结果显示,MARSE可实现SE性能与计算成本之间的灵活权衡。音频示例和代码可在线获取。

英文摘要

Most previous work on speech enhancement (SE) based on masked generative modeling relied on discrete token representations of audio signals, obtained using neural audio codecs (NACs). However, a recent study has shown that continuous latent representations of NACs can be advantageous for SE in terms of speech quality and intelligibility. In this work, we propose masked autoregressive SE (MARSE), a method for SE based on iterative decoding of masked clean speech frames using continuous NAC representations of speech. In particular, we investigate a set of different decoding policies, ceteris paribus, that is, using the same DNN (a Conformer model), the same NAC (the DAC codec) and the same training setup. The results show that MARSE enables a flexible trade-off between SE performance and computational cost. Audio examples and code are available online.

发表机构

  • CentraleSupélec(中央高等电力学院)
  • Univ. Grenoble Alpes(格勒诺布尔阿尔卑斯大学)
  • CNRS(法国国家科学研究中心)
  • Grenoble-INP(格勒诺布尔综合理工学院)
  • GIPSA-lab(吉普萨实验室)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑