arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

低资源多语言文本到语音的复杂文本鲁棒性评估与失败诊断

Complex-Text Robustness Evaluation and Failure Diagnosis for Low-Resource Multilingual Text-to-Speech

Tianlun Zuo, Ziyu Zhang, Tingzhi Mao, Zhonghua Fu, Lei Xie

arXiv 2609.11545首次发表:更新:

发表机构

Northwestern Polytechnical University; iFLYTEK Company Ltd.(西北工业大学; 科大讯飞股份有限公司)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出针对低资源多语言TTS的复杂文本鲁棒性诊断框架,从内容、语言和生成稳定性三维度评估,并引入轻量级文本风险评分(TRS),在多个系统上验证了其有效性。

AI 中文摘要

低资源多语言文本到语音(TTS)系统扩大了语言覆盖范围,但其在复杂文本输入下的鲁棒性仍未得到充分诊断。现有评估主要使用常规测试句子,关注自然度、说话人相似性和内容一致性,而对多语言TTS系统在处理数字、日期、命名实体、长句、语码转换表达及标点相关结构等挑战性输入时的失败方式提供的洞察有限。本文提出了一种针对低资源多语言TTS的复杂文本鲁棒性诊断框架。我们从三个维度评估鲁棒性:内容一致性、语言一致性和生成稳定性。为泰语、越南语、斯瓦希里语和印尼语设计了一个多语言鲁棒性测试方案,涵盖普通句子和多种类型的复杂文本输入。我们进一步引入了自动诊断指标,包括字符错误率、语言识别准确率和时长异常率。为了在语音生成前支持输入级风险分析,我们提出了一种轻量级文本风险评分(TRS),该评分从可解释的文本特征中估计合成风险,无需人工标注或模型训练。在包括OmniVoice、VoxCPM2和MMS-TTS在内的三个代表性多语言TTS系统上的实验表明,复杂文本输入暴露了常规短句评估未完全反映的系统性失败模式。不同系统在数字归一化、命名实体处理、长文本生成和语码转换输入处理方面表现出不同的脆弱性。此外,TRS与内容错误和时长异常呈正相关,证明了其作为低资源多语言TTS中复杂文本风险诊断的低成本合成前指标的有效性。

英文摘要

Low-resource multilingual text-to-speech (TTS) systems have expanded language coverage, but their robustness under complex text inputs remains insufficiently diagnosed. Existing evaluations mainly focus on naturalness, speaker similarity, and content consistency using regular test sentences, while providing limited insight into how multilingual TTS systems fail when handling challenging inputs such as numbers, dates, named entities, long sentences, code-switched expressions, and punctuation-related structures. This paper proposes a complex-text robustness diagnosis framework for low-resource multilingual TTS. We evaluate robustness from three dimensions: content consistency, language consistency, and generation stability. A multilingual robustness testing scheme is designed for Thai, Vietnamese, Swahili, and Indonesian, covering ordinary sentences and multiple types of complex text inputs. We further introduce automatic diagnostic metrics, including character error rate, language identification accuracy, and duration abnormal rate. To support input-level risk analysis before speech generation, we propose a lightweight Text Risk Score (TRS), which estimates synthesis risk from interpretable text features without manual annotation or model training. Experiments on three representative multilingual TTS systems, including OmniVoice, VoxCPM2, and MMS-TTS, show that complex text inputs expose systematic failure patterns that are not fully reflected by ordinary short-sentence evaluation. Different systems exhibit distinct vulnerabilities in number normalization, named entity handling, long-text generation, and code-switched input processing. Furthermore, TRS shows a positive correlation with content errors and duration abnormalities, demonstrating its usefulness as a low-cost pre-synthesis indicator for complex-text risk diagnosis in low-resource multilingual TTS.

CommentsNCMMSC 2026 accepted

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑