采样之前先看结构:面向数据高效文本转语音的社区感知核心集选择
Structure Before Sampling: Community-Aware Core-Set Selection for Data-Efficient Text-to-Speech
- Brac University(布拉克大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本文提出社区代表选择器,利用音位配位图结构挑选核心集,在孟加拉语和英语中提升稀有音素覆盖并降低TTS字符错误率,同时减少训练时间。
AI中文摘要:
文本转语音(TTS)语料库的录制成本高昂,但许多话语增加的语音学新信息很少。核心集选择通过在固定音频时长预算下挑选小型训练子集来降低这一成本。我们将语料库表示为音位配列图,该图将每个话语与其音素上最相似的话语相连,并首先测试该图是否具有结构。在孟加拉语和英语语料库中,其聚类系数分别是规模匹配的随机图的199倍和56倍,其模块度是保持度序列的随机图的两倍以上。随后我们提出社区代表(Community Representative)选择器,它跨图社区采样并在每个社区内分散选择,从富含稀有音素的话语开始。在每种预算和两种语言下,它覆盖的稀有音素双连字都比随机和基于熵的选择更多,且这一优势在保留话语上依然成立。在其20%核心集上训练的TTS模型,在两种语言中的字符错误率(CER)均显著低于在等时长随机或基于熵的子集上训练的模型。当所有模型训练相同轮数时,孟加拉语核心集模型还优于全语料库训练(CER为3.93%对4.47%),且训练时间减少4.5倍。
英文摘要:
Text-to-speech (TTS) corpora are costly to record, yet many utterances add little new phonetic information. Core-set selection reduces this cost by choosing a small training subset under a fixed audio-duration budget. We represent a corpus as a phonotactic graph that links each utterance to its most phonemically similar ones, and we first test whether this graph has structure. In Bangla and English corpora, its clustering is 199 and 56 times that of a size-matched random graph, and its modularity is more than twice that of a degree-preserving random graph. We then propose Community Representative, a selector that samples across graph communities and spreads its choices within each one, starting from utterances rich in rare phonemes. At every budget and in both languages, it covers more rare phoneme bigrams than random and entropy-based selection, and this lead holds on held-out utterances. TTS models trained on its 20% core-sets have a significantly lower character error rate (CER) than models trained on equal-duration random or entropy-based subsets in both languages. When all models train for the same number of epochs, the Bangla core-set model also outperforms full-corpus training (3.93% vs. 4.47% CER) with 4.5x less training time.