arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

预算内的声调:面向大规模多语言文本转语音的无参考词汇声调度量

Tone on a Budget: A Reference-Free Metric for Lexical Tone in Massively Multilingual Text-to-Speech

Moses Daudu, Adeola Enitan Bamidele, Honor-Jesus Bezaleel

arXiv 2609.14817首次发表:更新:

发表机构

Landmark University; Federal University of Agriculture, Abeokuta(地标大学; 联邦农业大学阿贝奥库塔分校)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对TTS中声调评估缺失的问题,提出无参考词汇声调度量DunDun,利用输入文本变音符号和音频音高轨迹,验证其有效性并揭示CER的局限。

AI 中文摘要

在约鲁巴语中,仅凭音高就能区分\d{o}k\d{o}(丈夫,中调)、\d{o}k\d{ò}(车辆,低调)和\d{o}k\d{ó}(锄头,高调)——这些变音符号本身就是声调标记。然而,字符错误率(CER)作为文本转语音(TTS)的标准自动化度量,在实践中是根据ASR输出计算的,而ASR输出会丢弃这些标记:一个合成器可能在CER上表现优异,却把“丈夫”说成“车辆”。我们提出了DunDun——以约鲁巴语中仅凭音高说话的说话鼓dùndún命名——一种自动化的、无参考的词汇声调度量,它不需要带声调标注的语料库。金标准的高/中/低序列从输入文本的变音符号中读取(在TTS中,该文本天然存在,因此无需参考录音);预测来自音频的音高轨迹。我们通过三种方式验证。使用PSOLA重合成展平音高会使DunDun崩溃,而CER不变。在300条母语者录音的答案键中反转高调和低调,使二分类读数降至0.14,对称地低于其0.35的随机水平——这是对评分路径的一致性检查,而非独立证据。三位母语听众在67次盲测A/B试验中,89.6%的时间选择了声调正确的片段(95%置信区间80.0-94.8;p<1e-4);在此样本量下,DunDun是否逐次跟踪这些判断尚未解决。应用于大规模多语言零样本TTS模型时,DunDun展示了CER无法显示的内容:在任何约鲁巴语微调之前,约鲁巴语声调接近母语锚点(五个解码种子的0.567±0.02对比0.596;随机水平0.33),尽管该模型自己的论文报告了21.4%的CER;而几小时的干净音频将CER减半(5小时从5.6%降至2.7%,15小时降至1.7%),同时声调在一小时内饱和。在非声调语言斯瓦希里语上,CER已经捕捉到了改进:语言所需的度量取决于语言本身。我们发布了该度量及完整的验证协议。

英文摘要

In Yorùbá, pitch alone separates \d{o}k\d{o} (husband, Mid), \d{o}k\d{ò} (vehicle, Low), and \d{o}k\d{ó} (hoe, High) -- the diacritics ARE the tone marks. Yet character error rate (CER), the standard automated metric for text-to-speech (TTS), is in practice computed from ASR output that drops those marks: a synthesizer can ace CER and still say vehicle for husband. We introduce DunDun -- named for the dùndún, the Yorùbá talking drum that speaks through pitch alone -- an automated, reference-free lexical-tone metric that needs no tone-labelled corpus. The gold High/Mid/Low sequence is read from the input text's diacritics (in TTS that text exists by construction, so no reference recording is needed); the prediction comes from the audio's pitch track. We validate three ways. Flattening pitch with PSOLA resynthesis collapses DunDun while CER does not move. Inverting High and Low in the answer key of 300 native recordings drives the two-class readout to 0.14, symmetrically below its 0.35 chance level -- a consistency check on the scoring path, not independent evidence. And three native listeners, over 67 blind A/B trials, pick the tone-correct clip 89.6% of the time (95% CI 80.0-94.8; p < 1e-4); whether DunDun tracks those judgements trial by trial is not resolved at this sample size. Applied to a massively multilingual zero-shot TTS model, DunDun shows what CER cannot: Yorùbá tone sits near the native anchor before any Yorùbá fine-tuning (0.567 +/- 0.02 over five decode seeds vs. 0.596; chance 0.33), despite the 21.4% CER the model's own paper reports; and a few hours of clean audio halve CER (5.6% to 2.7% by 5h, 1.7% by 15h) while tone saturates within the hour. On non-tonal Swahili, CER already captures the gains: the metric a language needs is language-dependent. We release the metric and the complete validation protocol.

Comments9 pages, 2 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑