arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.25351eess.AScs.SD

通过基于梯度的逆优化从冻结的文本转语音模型中提取语音风格

Extracting Voice Styles from Frozen TTS Models via Gradient-Based Inverse Optimization

  • Hanyang University(汉阳大学)

机构由 AI 辅助整理,请以论文原文为准。

Gyeongmin Kim

AI总结:

研究针对文本到语音系统中用户无法生成自身语音风格向量的问题,通过梯度下降直接求解输入,冻结模型权重仅优化风格向量,在多说话者实验中提升了相似度,验证器接受率显著提高。

AI中文摘要:

一些文本到语音系统提供合成模型和预设风格向量,但没有将音频转换为这种向量的参考编码器。模型仍接受风格向量,用户无法生成自己的风格向量。我们通过梯度下降直接求解该输入,反转已发布的管道:所有权重保持冻结,仅针对一个录音的时间池化WavLM统计信息优化风格向量。由于目标丢弃时间轴,合成文本可能与录音不同,因此不需要转录本和对齐。在来自两个语料库的154个说话者上,ECAPA - TDNN相似度从0.132提高到0.413,ResNet从0.099提高到0.401,每个说话者都有改善;在等误码点的验证器接受53%的恢复语音作为目标,而从预设语音开始的接受率为1%。

英文摘要:

Some text-to-speech systems ship a synthesis model and preset style vectors but withhold the reference encoder that turns a recording into a style vector, so a user cannot obtain a style for a new voice. We recover that vector without the encoder by inverting the released pipeline with gradient descent: all weights stay frozen and only the style vector is optimized, against time-pooled WavLM statistics of one recording of the target. The objective discards the time axis, so no transcript is needed. Two constraints from the release make the result a drop-in style: every row of the timbre style stays on the unit sphere where the presets lie, and the small duration style, which a time-pooled loss cannot see, is fitted to the recording's speaking rate. On SupertonicTTS, over 147 speakers and 100 held-out sentences each, ECAPA-TDNN similarity rises from 0.129 to 0.419, every recovered style starts from a preset and ends closer to its target than that preset was, and the pooled word error rate stays below that of the presets. A verifier at its equal-error point accepts 52% of the recovered voices as the target speaker, against 1% of the presets.

补充信息

↑