arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

西班牙语语音的客观可懂度度量评估

An Objective Intelligibility Metric Evaluation on Spanish Speech

Iván López-Espejo, Jesper Jensen

arXiv 2607.10619首次发表:更新:

AI 中文总结

研究在西班牙语语音可懂度数据集SpInt上评估五个基于参考的和两个基于深度学习的无参考客观可懂度度量,发现基于参考的度量表现更优,因开发时未接触西班牙语数据,无参考方法受训练测试声学不匹配影响大,故发布SpInt以推动相关研究。

AI 中文摘要

客观可懂度度量(OIM)能够快速且低成本地评估语音可懂度,在语音技术评估中广泛应用。本研究在新的西班牙语语音可懂度数据集SpInt上评估了五个基于参考的OIM(STOI、ESTOI、STGI、HASPI和SIIB)以及两个基于深度学习的无参考度量(MOSA-Net+和W2V-SIP)。结果表明基于参考的OIM始终优于现代数据驱动的无参考方法。由于评估指标在开发期间未接触西班牙语语音数据,这种效应在该场景中尤为显著。因此,为促进对更强大、更通用的无参考OIM的研究,SpInt被公开发布。

英文摘要

Objective intelligibility metrics (OIMs) enable fast and low-cost evaluation of speech intelligibility and are widely used in speech technology assessment. In this study, we evaluate five reference-based OIMs (STOI, ESTOI, STGI, HASPI, and SIIB) and two deep learning-based no-reference metrics (MOSA-Net+ and W2V-SIP) on SpInt, a new Spanish speech intelligibility dataset. Our results show that reference-based OIMs consistently outperform modern data-driven no-reference approaches, which degrade notably under training-test acoustic mismatches such as language mismatch. This effect is particularly relevant in our scenario, as none of the evaluated metrics were exposed to Spanish speech data during development. Consequently, to foster research on more robust and generalizable no-reference OIMs, SpInt is released publicly.

CommentsSubmitted to IberSPEECH 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑