arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

一种用于自适应决策导向语音增强的混合经典学习框架

A Hybrid Classical-Learning Framework for Adaptive Decision Directed Speech Enhancement

Ali Rajabi, Xiangwei Zhou

arXiv 2609.26183首次发表:更新:

AI 中文总结

提出一种混合经典学习框架ABCDD,通过帧相关下限增益界扩展决策导向方法,并用轻量级MLP预测beta参数,在VoiceBank-DEMAND上提升平均SNR 4.41 dB。

AI 中文摘要

语音增强旨在从带噪观测中恢复干净的语音信号,同时保持语音质量和可懂度。经典方法如谱减法和决策导向(DD)增强因其可解释性和低计算复杂度而仍被广泛使用,但在低信噪比(SNR)条件下,它们可能遭受音乐噪声伪影或对弱语音成分的过度衰减。本文提出了一种自适应Beta约束决策导向(ABCDD)语音增强框架,该框架通过帧相关的下限增益界扩展了传统的DD方法。引入的beta参数控制噪声抑制与语音保留之间的权衡。为了对大规模和多样化的数据集实现参数选择的自动化,进一步开发了一个轻量级多层感知器(MLP)模型,直接从带噪语音特征预测帧级beta值。所提出的框架通过一个代表性语音示例和在VoiceBank-DEMAND数据集上的大规模测试进行评估。在代表性示例中,ABCDD在多个客观指标上优于传统的谱减法和经典DD,包括SNR、对数谱距离(LSD)、均方根误差(RMSE)、相关性和尺度不变信号失真比(SI-SDR)。在100个未见过的VoiceBank-DEMAND测试文件上,所提出的MLP-beta ABCDD方法将平均尺度对齐SNR从9.41 dB提高到13.82 dB,对应平均增益为4.41 dB。结果表明,将可解释的经典增强结构与轻量级基于机器学习的参数自适应相结合,为鲁棒语音增强提供了一个有效且实用的方向。

英文摘要

Speech enhancement aims to recover clean speech signals from noisy observations while preserving speech quality and intelligibility. Classical methods such as Spectral Subtraction and Decision-Directed (DD) enhancement remain widely used because of their interpretability and low computational complexity, but they may suffer from musical-noise artifacts or excessive attenuation of weak speech components under low signal-to-noise ratio (SNR) conditions. This paper proposes an Adaptive Beta-Constrained Decision-Directed (ABCDD) speech enhancement framework that extends the conventional DD method through a frame-dependent lower gain bound. The introduced beta parameter controls the tradeoff between noise suppression and speech preservation. To automate parameter selection for large and diverse datasets, a lightweight multilayer perceptron (MLP) model is further developed to predict frame-level beta values directly from noisy-speech features. The proposed framework is evaluated using both a representative speech example and large-scale testing on the VoiceBank-DEMAND dataset. In the representative example, ABCDD outperformed conventional Spectral Subtraction and classical DD across multiple objective metrics, including SNR, Log-Spectral Distance (LSD), Root-Mean-Square Error (RMSE), correlation, and Scale-Invariant Signal-to-Distortion Ratio (SI-SDR). On 100 unseen VoiceBank-DEMAND test files, the proposed MLP-beta ABCDD method improved average scale-aligned SNR from 9.41 dB to 13.82 dB, corresponding to an average gain of 4.41 dB. The results indicate that combining interpretable classical enhancement structure with lightweight machine-learning-based parameter adaptation provides an effective and practical direction for robust speech enhancement.

CommentsNovelty is not clear

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑