arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.39088cs.SDcs.AI

SCIC:面向说话人自适应表现力TTS的范围与码本感知指令条件化

SCIC: Scope- and Codebook-Aware Instruction Conditioning for Speaker-Adapted Expressive TTS

Longyu Lu, Zongwei Du, Mengtao Xing, Zhuoqun Liu, Zifan Guan, Meiguang Jin, Junfeng Ma

AI总结:

针对长直播TTS缺乏从句级韵律控制的问题,提出SCIC方法,结合时间指令路由与码本加权,并采用多奖励GDPO后训练,提升说话人相对音高能量控制及段落表现力层次。

AI中文摘要:

长篇幅直播式TTS需要上下文相关的韵律和段落级连贯性。然而,许多现有的基于指令的TTS系统采用全局或统一条件,对从句级相对韵律变化的显式控制有限。我们引入了说话人相对内联韵律控制,其中每个音高、能量或速度指令都针对相对于同一说话人前一个从句的从句,而暂停则使用绝对持续时间间隔。在基于编解码器的TTS中,速度和暂停影响序列长度,而音高和能量依赖于残差码本。通过分析Qwen3-TTS RVQ码本,我们发现能量集中在早期残差码本中,而音高则累积在更深的序列前缀中。因此,我们提出了范围与码本感知指令条件化(SCIC),将用于帧级标签激活的时间指令路由器与残差码本上的标签特定码本加权相结合。SCIC在标准指令微调使用文本令牌标签的基础上,改善了说话人相对音高和能量控制。我们进一步应用多奖励GDPO后训练来联合优化控制和质量,在保持CER和说话人相似性的同时提高了控制准确性。在长篇幅合成中,SCIC比没有指令的说话人自适应SFT产生了更明显的段落级表现力层次。音频演示可在以下网址获取:this https URL

英文摘要:

Long-form live-streaming TTS requires context-dependent prosody and paragraph-level coherence. However, many existing instruction-based TTS systems use global or uniform conditions, providing limited explicit control over clause-level relative prosodic changes. We introduce Speaker-Relative Inline Prosody Control, where each Pitch, Energy, or Speed instruction targets a clause relative to the preceding clause from the same speaker, while Pause uses an absolute duration interval. In codec-based TTS, Speed and Pause affect sequence length, whereas Pitch and Energy rely on residual codebooks. By analyzing Qwen3-TTS RVQ codebooks, we find that Energy concentrates in early residual codebooks, whereas Pitch accumulates across a deeper prefix. We therefore propose Scope- and Codebook-Aware Instruction Conditioning (SCIC), combining a Temporal Instruction Router for frame-level tag activation with Tag-Specific Codebook Weighting over residual codebooks. SCIC improves speaker-relative Pitch and Energy control over standard instruction fine-tuning using text-token tags. We further apply multi-reward GDPO post-training to jointly optimize control and quality, improving control accuracy while preserving CER and speaker similarity. In long-form synthesis, SCIC produces a more distinct paragraph-level expressive hierarchy than speaker-adapted SFT without instructions. Audio demos are available at: https://taoliveaigc.github.io/SCIC/

↑