arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.34662cs.SDeess.AS

基于漂移的无监督语音增强

Unsupervised Speech Enhancement via Drifting

Diego Caviedes-Nozal, Liang Xu, Rasmus Kongsgaard Olsson, W. Bastiaan Kleijn

首次发表
浏览论文内容

中文总结 AI 辅助

本文提出输入条件化漂移方法,解决无监督语音增强中漂移导致的内容和说话人身份丢失问题,通过锚定编码器和键编码器重绑输入,显著降低词错误率并恢复说话人相似度。

中文摘要 AI 辅助

本文研究了在无配对设置下使用漂移方法进行无监督语音增强,其中训练依赖于分别收集的退化音频和干净音频,而不需要对应的配对。虽然最近的漂移方法能够实现无配对训练,但代价高昂:由于目标函数仅优化干净语音的边缘先验,增强器会逐渐丢失输入的语言内容和说话人身份。为解决这一问题,我们引入了输入条件化漂移。我们在保持干净语料库吸引力的同时,通过两种机制将输出重新绑定到退化输入:锚定编码器通过向输入特征靠拢来提供缺失的似然,键编码器通过对检索到的帧进行重新加权来调节先验。两者都不需要标签或配对数据。使用无需训练的编码器选择标准,VoiceBank-DEMAND上的词错误率降至10.1%(未处理时为11.7%),说话人相似度从0.490恢复到0.879,该方案部分迁移至WSJ0-REVERB上的去混响任务:内容得到改善,但渲染质量未提升。

英文摘要

This paper addresses unsupervised speech enhancement in the unpaired setting using drifting methods, where training relies on separate collections of degraded and clean audio without corresponding pairs. While recent drifting approaches enable unpaired training, they do so at a heavy cost: because the objective optimizes only a marginal prior over clean speech, the enhancer gradually loses the input's linguistic content and speaker identity. To fix this, we introduce input-conditioned drifting. We preserve the pull of the clean corpus while re-tethering the output to the degraded input via two mechanisms: an anchor encoder supplies the missing likelihood by pulling toward the input's features, and a key encoder conditions the prior by re-weighting retrieved frames. Neither requires labels or paired data. Using a training-free encoder selection criterion, Word Error Rate on VoiceBank-DEMAND falls to 10.1% (unprocessed: 11.7%), speaker similarity recovers from 0.490 to 0.879, and the recipe transfers in part to dereverberation on WSJ0-REVERB: content improves, rendering quality does not.

发表机构

  • GN A/S(GN 公司)
  • Victoria University of Wellington(惠灵顿维多利亚大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑