发表机构
FAST School of Computing, National University of Computer and Emerging Sciences (FAST-NUCES)(FAST计算机学院,国家计算机与新兴科学大学(FAST-NUCES))
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出可复现的多指标TTS评估框架,针对低资源语言的四类语音领域评估四款SOTA TTS系统,发现情感语音合成挑战最大、对话语音保真度最高,公开资源支持相关研究。
AI 中文摘要
近期神经文本语音合成(TTS)系统的进展已大幅提升了多种语言的语音自然度与可懂度,但针对不同语音领域联合评估感知质量、说话人相似度与声学保真度的综合评估方法仍较为有限,尤其是针对低资源和代表性不足的语言。本文提出一种可复现的多指标基准框架,用于通过领域特定分析系统评估现代TTS系统;该框架整合了互补的主观与客观评估协议,并通过针对某代表性低资源语言的综合案例研究进行验证,覆盖正式、对话、文学/讲故事、情感四类语音领域。研究采用MUSHRA听觉测试、ABX区分测试、基于Resemblyzer的说话人相似度评分,以及基于梅尔倒谱失真(MCD)和基频均方根误差(F0 RMSE)的声学分析,对Indic-Parler-TTS、MMS-TTS、Microsoft Edge TTS、Google Gemini TTS四款最先进TTS系统进行评估,共涉及960组音频对。结果显示,TTS系统在不同语音领域的性能存在显著差异,其中情感语音始终是合成挑战最大的领域(平均MCD为12.03 dB,平均F0 RMSE为889 cents),而对话语音的整体声学保真度最高。除实证发现外,本研究还提供了可复现的评估框架,公开了评估脚本、结果表格和可执行的Colab笔记本,以支持标准化基准测试及未来针对低资源语言TTS评估的研究。
英文摘要
Recent advances in neural text-to-speech (TTS) systems have substantially improved speech naturalness and intelligibility across many languages. However, comprehensive evaluation methodologies that jointly assess perceptual quality, speaker similarity, and acoustic fidelity across diverse speech domains remain limited, particularly for low-resource and underrepresented languages. This paper presents a reproducible, multi-metric benchmarking framework for systematic evaluation of modern TTS systems through domain-specific analysis. The proposed framework integrates complementary subjective and objective evaluation protocols and is demonstrated through a comprehensive case study on a representative low-resource language spanning four speech domains: Formal, Conversational, Literary/Storytelling, and Emotional. Four state-of-the-art TTS systems -- Indic-Parler-TTS, MMS-TTS, Microsoft Edge TTS, and Google Gemini TTS -- are evaluated using MUSHRA listening tests, ABX discrimination tests, speaker similarity scoring with Resemblyzer, and acoustic analyses based on mel-cepstral distortion (MCD) and F0 RMSE over 960 audio pairs. Results reveal substantial variation in TTS performance across speech domains, with emotional speech consistently presenting the greatest synthesis challenge (mean MCD 12.03 dB; mean F0 RMSE 889 cents), while conversational speech achieves the highest overall acoustic fidelity. Beyond the empirical findings, this work provides a reproducible evaluation framework, publicly releasing evaluation scripts, result tables, and executable Colab notebooks to support standardized benchmarking and future research on TTS evaluation for low-resource languages.
Comments17 pages, 1 figure. Submitted to Computer Speech & Language (Elsevier)