arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.07781eess.AS

通过推理时的重新思考与修正模块缓解语音增强中的过度抑制问题

Mitigating Over-Suppression in Speech Enhancement via Inference-Time Rethink-and-Refine Correction Module

Mike Qu, Yu-Wen Chen, Julia Hirschberg

AI总结:

该研究提出推理阶段无需额外训练的rethink-and-refine修正模块,通过识别增强不可靠区间并选择性重新混合,在多数据集上提升语音增强的感知质量、可懂度及下游性能。

AI中文摘要:

我们提出了一种重新思考与修正(rethink-and-refine)模块,用于解决语音增强(SE)模型的常见失效模式——过度抑制,即语音线索与噪声一同被抑制的问题。该方法完全在推理阶段运行,无需额外训练,可无缝集成到各类SE模型中。给定带噪信号与增强后信号,我们利用自动语音识别模型获取词级或音素级对齐结果,识别增强不可靠的区间,随后通过凸插值对这些区间进行选择性重新混合,其中每段权重经优化以最大化平衡感知质量与语音保留度的复合目标。在URGENT 2024、URGENT 2025、VCTK-DEMAND及MSP-PODCAST数据集上的实验表明,与单独使用传统SE方法相比,该方法在感知质量、可懂度及下游性能上均实现了一致提升,证明了重新思考与修正框架对稳健语音处理的价值。

英文摘要:

We present a rethink-and-refine correction module that addresses over-suppression, a common failure mode of speech enhancement (SE) models, where speech cues are suppressed alongside noise. Our method operates entirely in the inference stage without additional training, allowing seamless integration with diverse SE models. Given noisy and enhanced signals, we obtain word- or phoneme-level alignments using an automatic speech recognition model and identify intervals where enhancement is unreliable. These intervals are then selectively remixed through convex interpolation, with per-segment weights optimized to maximize a composite objective balancing perceptual quality and speech preservation. Experiments on the URGENT 2024 and 2025, VCTK-DEMAND, and MSP-PODCAST datasets show consistent improvements in perceptual quality, intelligibility, and downstream performance compared to conventional SE alone, demonstrating the benefit of rethink-and-refine framework for robust speech processing.

补充信息

↑