通过激活引导控制自回归TTS中的语速
Controlling Speaking Rate in Autoregressive TTS via Activation Steering
浏览论文内容
中文总结 AI 辅助
本文提出一种无需重训练的自回归TTS语速控制方法,通过钳制解码器激活沿语速轴的投影实现推理时语速引导,保持说话人身份与自然度,并在Seed-TTS-Eval基准上验证有效性。
中文摘要 AI 辅助
自回归文本转语音(TTS)系统能够合成自然语音,但一旦训练完成,对语速的控制能力十分有限。我们证明,在推理时无需重新训练,仅通过将单个解码器块的激活沿一条发现的语速轴进行钳制,即可引导语速。对解码器块的分析可恢复语速轴、中性工作点以及每步的强度标度;在推理时,激活在该轴上的投影被设置为一个固定标量。从合成的时间拉伸和时间压缩语音中学习这一方向,能够实现语速控制,且在很大程度上保持说话人身份,跨模型架构泛化,并在客观和人工评估中保持高自然度。与在慢速极端情况下失效的标准加法引导不同,钳制在所测试的三个系统上均保持稳定;在中等目标下,更优的规则取决于模型。最后,我们证明语速信息在各层中均可解码,但仅在中间深度窗口内可因果引导,并在公开的Seed-TTS-Eval基准上展示了我们方法的有效性。
英文摘要
Autoregressive text-to-speech (TTS) systems synthesize natural speech but, once trained, offer little control over speaking rate. We show that speaking rate can be steered at inference time, without retraining, by clamping a single decoder block's activation along a discovered speed axis. A decoder-block analysis recovers the rate axis, a neutral operating point, and a per-step intensity scale; at inference, the activation's projection onto this axis is set to a fixed scalar. Learning this direction from synthetically time-stretched and time-compressed speech yields rate control that largely preserves speaker identity, generalizes across model architectures, and maintains high naturalness in objective and human evaluations. Unlike standard additive steering, which breaks at the slow extreme, clamping remains stable on all three systems tested; at moderate targets, the better rule depends on the model. Finally, we show that rate information is decodable across layers but causally steerable only within a mid-depth window, and demonstrate the effectiveness of our approach on the public Seed-TTS-Eval benchmark.
发表机构
- AGIGO
- ETH Zurich(苏黎世联邦理工学院)
- Sapienza University of Rome(罗马大学)
- University of Zurich(苏黎世大学)
机构由 AI 辅助整理,请以论文原文为准。