arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

LIFT-SE:语言推断与流变换相结合的生成式语音增强

LIFT-SE: Linguistic Inference Followed by Flow Transformation for Generative Speech Enhancement

Haoyin Yan, Chengwei Liu, Zheng Xue, Xiaotao Liang, Jifa Cai, Zeyu Zhao, Jingjing Wang

arXiv 2610.09963首次发表:更新:

发表机构

Qwen Business Unit of Alibaba(阿里巴巴通义千问业务部)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

提出LIFT-SE两阶段生成框架,解耦语言推断与声学合成,结合离散令牌预测和连续流匹配,在混响下提升语言一致性并保持感知质量。

AI 中文摘要

在严重噪声和混响条件下,语义约束不可靠时,生成式语音增强(SE)容易产生语言幻觉。此外,生成离散编解码器令牌的方法受限于编解码器解码器的量化误差,无论令牌预测的准确性如何。我们提出LIFT-SE,一个两阶段生成框架,在QRes-Codec中将语言推断与声学合成解耦,该编解码器暴露了量化潜变量及其残差补全的连续形式。第一阶段自回归地预测干净的编解码器令牌,条件为来自自监督前端(向干净语音蒸馏)的帧对齐特征。第二阶段应用条件流匹配,将高斯噪声传输到以预测令牌为条件的连续潜变量,冻结的解码器重建增强波形。离散生成提供自然度,而连续细化恢复信号保真度。在DNS1和URGENT基准上的实验表明,LIFT-SE在混响条件下获得了良好的语言一致性以及有竞争力的感知质量,系统的消融实验验证了两个阶段的必要性。代码将在未来发布。

英文摘要

Generative speech enhancement (SE) is prone to linguistic hallucination when semantic constraint is unreliable under severe noise and reverberation. Moreover, approaches that generate discrete codec tokens are bounded by the quantization error of the codec decoder, regardless of token-prediction accuracy. We propose LIFT-SE, a two-stage generative framework that decouples linguistic inference from acoustic synthesis within QRes-Codec, which exposes a quantized latent and its residual-completed continuous form. The first stage predicts clean codec tokens autoregressively, conditioned on frame-aligned features from a self-supervised front-end distilled toward clean speech. The second stage applies conditional flow matching to transport Gaussian noise to the continuous latent conditioned on the predicted tokens, and the frozen decoder reconstructs the enhanced waveform. Discrete generation provides naturalness, while continuous refinement restores signal fidelity. Experiments on the DNS1 and URGENT benchmarks show that LIFT-SE attains favorable linguistic consistency under reverberant conditions together with competitive perceptual quality, and systematic ablations verify the necessity of both stages. Code will be released in the future.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑