发表机构
BitmanagerAI; lab260; MTUCI(BitmanagerAI; lab260; 莫斯科通信与信息技术大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究隔离了神经文本到语音模型在重复文本上的计数失败,证明重复本身而非长度是主因,并通过对照实验和多种分析验证了该现象的稳健性。
AI 中文摘要
文本到语音模型在文本多次重复某个短语时会出现循环、截断和计数丢失的问题。我们证明,导致模型失效的是重复本身,而非随之而来的长度。在我们的测试集中,每个重复句子都配有一个对照句,其句子数和单词数匹配,但没有任何单词会连续重复。来自三种架构的六个模型几乎完美地处理了对照句,却在重复的孪生句上失败:在k≥6时,完全正确的比例分别为94.3%和18.2%。这一差距在贪婪解码、重复惩罚扫描、四种独立的语音识别器以及420种分析规范下均未逆转符号;一个留出的第四种架构的预测差距误差在一个百分点以内,两个非自回归基线中有一个表现出相同的失败。改变文本的周期表明,失败随周期性平滑增长,当没有单词与自身相邻时,仍有半数失败保留。
英文摘要
Text-to-speech models loop, truncate and lose count on text that repeats a phrase many times. We show that repetition itself is what breaks them, not the length that comes with it. Every repeated sentence in our test set is paired with a control of matched sentence and word count in which no word ever repeats back-to-back. Six models from three architectures render the controls almost perfectly and fail the repeated twins: 94.3% against 18.2% exactly right at k >= 6. The gap survives greedy decoding, repetition-penalty sweeps, four independent speech recognisers and 420 analysis specifications without once reversing sign; a held-out fourth architecture lands within a point of its predicted gap, and one of two non-autoregressive baselines shows the same failure. Varying the period of the text shows the failure grows smoothly with periodicity, half of it surviving when no word is adjacent to itself.
CommentsSubmitted to IEEE ICASSP 2027. Code and data: https://github.com/lab260ru/tts-counting-failure