arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.31892cs.SDcs.AIcs.CLeess.AS

NVAlign:连续自回归流匹配文本到语音中非语言控制的直接梯度优化

NVAlign: Direct-Gradient Optimization for Non-Verbal Control in Continuous Autoregressive Flow Matching Text-to-Speech

Qiaolin Wang, Pedro Sandoval-Segura, Anunaya Joshi, Edvardas Jurkonis, Jake Downie

首次发表
浏览论文内容

中文总结 AI 辅助

NVAlign通过直接梯度优化,在连续自回归流匹配TTS中实现对非语言发声标签的精确控制,优于现有基线。

中文摘要 AI 辅助

虽然现代文本到语音(TTS)系统能够生成高度自然的语音并支持内联非语言发声(NVV)标签,但对这些事件的精确控制仍然具有挑战性。一个关键差距在于,缺乏针对连续自回归流匹配TTS中非语言控制的既定训练后方法。为此,我们提出了NVAlign,一种用于该架构中NVV标签遵循的直接梯度训练后框架。我们首先在带有NVV标注的语音上对TTS模型和NVV感知的自动语音识别(NV-ASR)模型进行监督微调(SFT),然后冻结NV-ASR模型作为训练后的奖励模型。一个两步梯度替代方案使得奖励能够通过流匹配采样器高效地反向传播,从而联合更新自回归主干和声学流头。保真度惩罚和参考速度正则化有助于保持说话人相似性和语音质量。NVV-SuperBench和人工听力评估的结果表明,NVAlign在标签遵循准确性上优于SFT和Flow-GRPO基线。这些发现表明,直接奖励梯度优化能够改善连续自回归流匹配TTS中的非语言控制。音频样本可在该https URL获取。

英文摘要

While modern text-to-speech (TTS) systems generate highly natural speech and support inline non-verbal vocalization (NVV) tags, accurate control over these events remains challenging. A key gap is the lack of established post-training methods for non-verbal control in continuous autoregressive flow-matching TTS. To this end, we present NVAlign, a direct-gradient post-training framework for NVV tag-following in this architecture. We first perform supervised fine-tuning (SFT) of TTS models and an NVV-aware automatic speech recognition (NV-ASR) model on NVV-annotated speech, then freeze the NV-ASR model to serve as the reward model for post-training. A two-step gradient surrogate enables efficient reward backpropagation through the flow-matching sampler to jointly update the autoregressive backbone and acoustic flow head. Fidelity penalties and reference-velocity regularization help preserve speaker similarity and speech quality. Results from NVV-SuperBench and human listening evaluations show that NVAlign improves tag-following accuracy over SFT and Flow-GRPO baselines. These findings demonstrate that direct reward-gradient optimization can improve non-verbal control in continuous autoregressive flow-matching TTS. Audio samples are available at https://nvalign.github.io/.

补充信息

↑