感知质量损失还是感知质量的损失?
Perceptual Quality Loss or Loss of Perceptual Quality?
浏览论文内容
中文总结 AI 辅助
本文通过客观与主观实验发现,语音增强模型训练中加入PESQ损失虽能提高PESQ分数,但未必改善听感,且PESQ主导复合指标,警示过度依赖PESQ并强调完整评估的重要性。
中文摘要 AI 辅助
当代深度语音增强(SE)模型在训练时,常常在损失函数中加入特定的辅助项,以提升其在感知指标上的表现。然而,在感知指标上获得更高的分数并不一定意味着听感体验的改善。通过客观和主观实验,我们评估了使用两种不同类型辅助PESQ损失项训练的语音增强模型的性能。在一系列标准指标上的数值评估表明,虽然针对PESQ优化的模型在测试集上自然获得更高的PESQ分数,但对于大多数其他指标,分数并没有显著变化。在某些情况下,PESQ损失甚至在不匹配的数据上导致更差的PESQ分数。一项正式的听音实验显示,在所有设置下,没有PESQ损失的模型通常比包含PESQ损失的模型更受青睐。最后,我们分析了PESQ在复合指标CSIG、CBAK和COVL中的相对重要性,发现PESQ在所有复合指标中占主导地位。我们的研究强调了过度依赖PESQ的风险,并强调了完整评估流程对语音增强的重要性。
英文摘要
Contemporary deep speech enhancement (SE) models are often trained with specific auxiliary terms in the loss function as a way to improve their performance in terms of perceptual metrics. Nevertheless, a higher score on a perceptual metric does not necessarily correlate with an improved listening experience. Through objective and subjective experiments, we assess the performance of SE models trained with two different types of auxiliary PESQ loss terms. The numerical evaluation on a suite of standard metrics suggests that, while models optimized for PESQ naturally obtain higher PESQ scores in the test set, for most other metrics the scores do not significantly change. In some cases, the PESQ loss even results in worse PESQ scores on mismatched data. A formal listening experiment reveals that the models without a PESQ loss were generally preferred over models that include it, across all settings. Finally, we analyze the relative importance of PESQ in the composite metrics CSIG, CBAK and COVL, and find that PESQ dominates all of them. Our study highlights the perils of over-reliance on PESQ and stresses the importance of a complete evaluation procedure for SE.
发表机构
- University of Hamburg(汉堡大学)
机构由 AI 辅助整理,请以论文原文为准。