arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

噪声差异下的多类语音分类

Multiclass Speech Classification Under Noise Disparity

Mahdi Amiri, Sayantan Biswas, Mingchi Hou, Pascal Frossard, Ina Kodrasi

arXiv 2610.03381首次发表:更新:

AI 中文总结

本研究提出多类交叉增强方案,消除噪声与类别标签关联,防止分类器利用噪声差异,在多类情感识别中有效缓解噪声影响,优于语音增强预处理。

AI 中文摘要

我们研究了多类语音分类任务中的噪声差异问题,并制定了一种策略,以防止分类器利用依赖于类别的噪声特征。基于我们先前在二分类方面的工作,我们引入了一种多类交叉增强方案,该方案使每个类别暴露于其他类别的噪声特征,从而消除个体噪声条件与类别标签之间的关联。我们将这种基于训练的方法与作为预处理策略的语音增强(SE)进行比较,后者旨在直接从输入中抑制与噪声相关的线索。在多类情感识别上的实验表明,交叉增强能有效缓解一系列信噪比下噪声差异的影响,而语音增强对模型性能产生不利影响。

英文摘要

We investigate noise disparity in multi-class speech classification tasks and develop a strategy to prevent classifiers from exploiting class-dependent noise characteristics. Building on our previous work for binary classification, we introduce a multi-class cross-augmentation scheme that exposes each class to the noise characteristics of the other classes, thereby removing the association between individual noise conditions and class labels. We compare this training-based approach with speech enhancement (SE) as a preprocessing strategy, which aims to suppress noise-related cues directly from the input. Experiments on multi-class emotion recognition show that cross-augmentation effectively mitigates the effect of noise disparity across a range of signal-to-noise-ratios, while SE has a detrimental effect on model performance.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑