预测MOS在复现人类对语音增强系统级偏好方面的可靠性如何?
How Reliable Are Predicted MOS for Reproducing Human System-Level Preferences in Speech Enhancement?
- Waseda University(早稻田大学)
- NTT, Inc.(日本电信电话株式会社)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本研究提出系统级偏好准确率(SPA)指标,直接评估预测MOS与人类评分在系统比较上的一致性,发现单一模型准确率差异大(9.4%-76.8%),集成改进有限,领域自适应在封闭条件有效,提示仅凭预测MOS可能导致不可靠结论。
AI中文摘要:
我们通过引入系统级偏好准确率(SPA),研究预测的平均意见得分(MOS)能否可靠地支持语音增强(SE)方法的系统级比较。尽管MOS预测模型被广泛用于评估SE系统,但其性能通常通过与人类评分的MOS的相关性来评估,这并不能保证在哪个系统更好上达成一致。SPA通过直接评估预测的MOS和人类评分的MOS是否产生相同的系统偏好来解决这一差距。利用SPA,我们系统地评估了三种设置:单一预测模型、集成学习和领域自适应。通过实验,SPA在单一预测模型之间差异显著,从9.4%到76.8%不等。即使是最好的模型,在约23%的系统比较中也与人类判断不一致。集成学习仅带来有限的改进,而领域自适应在封闭条件下往往能大幅提高SPA,但在更实际的开集条件下(目标系统和说话者均未知)仅带来适度的增益。这些结果表明,SPA可以揭示仅基于相关性的评估无法暴露的错误,且仅凭预测的MOS可能导致在实际SE系统比较中得出不可靠的结论。
英文摘要:
We investigate whether predicted Mean Opinion Scores (MOS) can reliably support system-level comparisons of speech enhancement (SE) methods by introducing system-level preference accuracy (SPA). Although MOS prediction models are widely used to evaluate SE systems, their performance is typically assessed by correlation with human-rated MOS, which does not guarantee agreement on which system is better. SPA addresses this gap by directly evaluating whether predicted and human-rated MOS yield the same system preferences. Using SPA, we systematically evaluate three settings: single prediction models, ensembling, and domain adaptation. Through experiments, SPA varies substantially across single prediction models, from 9.4% to 76.8%. Even the best model disagrees with human judgments in approximately 23% of system comparisons. Ensembling yields only limited improvement, while domain adaptation tends to substantially improve SPA in the closed condition but brings only modest gains in the more practical open condition, where neither the target systems nor the speakers are known. These results suggest that SPA can reveal errors correlation-based evaluation alone does not expose, and that predicted MOS alone can lead to unreliable conclusions in practical SE system comparison.