arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.03283cs.SD

语义是否足以用于语音平均意见得分预测?

Is Semantics Enough for Speech Mean Opinion Score Prediction?

Tianyu Lan, Yufei Shi, Yang Ai, Honghao Sun, Huipeng Du, Zhenhua Ling

首次发表
浏览论文内容

中文总结 AI 辅助

该研究针对自动语音MOS预测器过度依赖SSL模型忽略声学细节的问题,对比三类模型在BVCC及OOD数据集的表现,发现结合语义与细粒度声学建模的模型性能更优,指出需同时关注语义与声学保真度。

中文摘要 AI 辅助

平均意见得分(MOS)是评估合成语音自然度的黄金标准。然而,当前自动MOS预测器以优先考虑高级语义的自监督学习(SSL)模型为主,这可能会损害其捕捉关键声学细节的能力。在本文中,我们系统研究了三种范式的表示:自监督学习(SSL)、仅声学的神经音频编解码器(NAC)以及将语义集成到基于重建的架构中的统一神经音频编解码器(NAC)。在标准BVCC数据集和多个域外(OOD)数据集上的广泛基准测试表明,将语义理解与细粒度声学建模相结合的特征在语音质量评估中实现了更高的性能上限。最终,我们的发现表明仅靠语义是不够的;对语义内容和声学保真度的双重关注对于可靠的MOS预测至关重要。

英文摘要

Mean Opinion Score (MOS) is the gold standard for evaluating synthesized speech naturalness. However, current automatic MOS predictors are dominated by self-supervised learning (SSL) models that prioritize high-level semantics, potentially compromising their ability to capture critical acoustic details. In this paper, we systematically investigate representations from three paradigms: SSLs, acoustic-only neural audio codecs (NACs), and unified NACs that integrate semantics into reconstruction-based architectures. Extensive benchmarking on the standard BVCC and multiple out-of-domain (OOD) datasets demonstrates that features synergizing semantic understanding with fine-grained acoustic modeling achieve a higher performance upper bound in speech quality assessment. Ultimately, our findings highlight that semantics alone are not enough; a dual focus on semantic content and acoustic fidelity is essential for robust MOS prediction.

发表机构

  • National Engineering Research Center of Speech and Language Information Processing, University of Science and Technology of China(中国科学技术大学国家语音与语言信息处理工程研究中心)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑