arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

精炼提升可懂度,搜索赋予身份:测试时计算在掩码扩散文本到语音合成中的作用

Refinement Buys Intelligibility, Search Buys Identity: What Test-Time Compute Buys in Masked-Diffusion TTS

Nityanand Mathur, Hamees Sayed, Ayush Pratap Singh

arXiv 2610.03320首次发表:更新:

发表机构

Blackstar Inc.; Smallest AI(黑星公司; Smallest AI)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究通过训练15个掩码扩散TTS模型并扫描推理步骤,发现精炼主要提升可懂度而搜索更有效恢复说话人身份,二者针对不同瓶颈,应分别优化。

AI 中文摘要

用于文本到语音合成的扩散语言模型结合了两种计算形式:模型深度(参数)和精炼步骤(推理预算)。我们探究它们在能力上是否同等扩展。我们在2000小时语音上训练了15个掩码扩散编解码器TTS模型,其深度各异(19-133M参数,3个随机种子),并在推理时扫描精炼步骤T在[1,16]范围内的取值,通过ASR词错误率(可懂度)和说话人验证(身份)对174个保留说话人进行零样本合成评估。相对于测量下限,精炼步骤弥补了可懂度范围的86.2%,但仅弥补了身份范围的46.4%——这一1.86倍的差异在多种错误指标下均稳健。以3倍和6倍调度重新训练会减弱但不会逆转这一差距(从1.84降至1.36再降至1.23倍),因为可懂度随步骤增加而饱和,而身份则持续改善。最佳K搜索在精炼失效时恢复了说话人身份,在四个独立编码器上的胜率为64.6%-79.0%。深度和步骤不可互换:可分离的B(d)B(T)模型拟合显著优于替代模型(Delta AICc=+69.3)。分析表明,剩余身份缺陷的62%存在于编解码器而非生成器中。我们得出结论:精炼和深度针对不同的瓶颈,应分别优化。

英文摘要

Diffusion language models for text-to-speech combine two forms of computation: model depth (parameters) and refinement steps (inference budget). We ask whether they scale equally across capabilities. We train 15 masked-diffusion codec TTS models varying depth (19-133M parameters, 3 seeds) on 2,000 hours of speech and sweep refinement steps T in [1,16] at inference, measuring zero-shot synthesis via ASR word error rate (intelligibility) and speaker verification (identity) on 174 held-out speakers. Against measured floors, refinement closes 86.2% of the intelligibility range but only 46.4% of the identity range - a 1.86x asymmetry robust across multiple error metrics. Retraining at 3x and 6x schedule attenuates but does not reverse this gap (1.84 to 1.36 to 1.23x), because intelligibility saturates with steps while identity continues improving. Best-of-K search recovers speaker identity where refinement fails, with 64.6-79.0% win rates across four independent encoders. Depth and steps are not interchangeable: separable B(d)B(T) fits significantly better (Delta AICc=+69.3) than substitution models. Analysis shows 62% of remaining identity deficit lies in the codec, not the generator. We conclude that refinement and depth target different bottlenecks and should be optimized separately.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑