arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.11593cs.SDeess.AS

Luna-TTS 系列技术报告

Luna-TTS Family Technical Report

Feng Yin, Shuai Shi, Junjie Zheng, Kechenying Zhou, Yiqiu Wang, Chenyang He, Qiuhua Jiang, Mengxiao Bi, Yanmin Qian, Mingxin Chen, Xun Gong, Tianteng Gu, Bing H… 展开作者

Feng Yin, Shuai Shi, Junjie Zheng, Kechenying Zhou, Yiqiu Wang, Chenyang He, Qiuhua Jiang, Mengxiao Bi, Yanmin Qian, Mingxin Chen, Xun Gong, Tianteng Gu, Bing Han, Peng Jiang, Chenda Li, Haiyang Sun, Han Wang, Wei Wang, Yi Wang, Leying Zhang, Wangyou Zhang, Chushu Zhou

首次发表
浏览论文内容

中文总结 AI 辅助

该研究提出基于扩散语言模型的 Luna-TTS 系列 TTS 系统,含非自回归与实时自回归变体,在多数据集上的语音识别、相似度及情感控制等指标优于对比系统。

中文摘要 AI 辅助

现代文本到语音(TTS)技术主要由自回归(AR)编解码器语言模型主导,其从左到右的解码会带来随话语长度增长的延迟、已生成前缀的误差累积,以及对残差向量量化(RVQ)标记网格施加的人工生成顺序。我们提出 Luna-TTS 系列,这是基于扩散语言模型的 TTS 系统,在涵盖中文、英文、日文、韩文的 100 万小时语音数据上进行预训练。该系列通过对预训练的 AR 文本大语言模型(LLM)进行逐步适配构建,从因果注意力到双向注意力,最终到块因果注意力,包含两个变体,共享同一个分词器、数据流水线和 0.6B 主干谱系。Luna-TTS 完全是非自回归的:它在固定数量的并行细化步骤中生成整个 RVQ 标记网格,零样本语音克隆和语音编辑作为填充功能自然实现。通过持续训练衍生的 Luna-TTS Realtime 是在 32 个编解码器帧(1.28 秒)块上自回归,同时并行去噪每个块;它支持键值(KV)缓存的块级生成和增量音频交付,在预热服务协议下实现端到端实时因子(RTF)为 0.0240,本地首块延迟为 41.6 毫秒。退火微调阶段增加了对情感和非语言发声(NVV)的显式控制,强化学习阶段应用 GRPO,其策略比率基于已实现的去噪轨迹计算。在 Seed-TTS-Eval 上,Luna-TTS 在所有四个指标上优于对比的开源和商业系统(中文测试集 CER 为 0.73、SIM 为 79.7,英文测试集 WER 为 1.49、SIM 为 76.8);在更具挑战性的野外 CV3-Eval 上,它在对比中实现了最低的普通话和英文错误率。与领先商业系统相比,它在 NVV 和情感控制的多数客观、基于模型及人类评估指标上取得最佳结果。

英文摘要

Modern text-to-speech (TTS) is dominated by autoregressive (AR) codec language models, whose left-to-right decoding brings latency that grows with utterance length, error accumulation along the committed prefix, and an artificial generation order imposed on the Residual Vector Quantization (RVQ) token grid. We propose Luna-TTS Family, diffusion-language-model-based TTS systems pretrained on 1 million hours of speech across Chinese, English, Japanese, and Korean. The family is built by progressive adaptation of a pretrained AR text LLM, from causal to bidirectional and finally to block-causal attention, and comprises two variants sharing a single tokenizer, data pipeline, and 0.6B backbone lineage. Luna-TTS is fully non-autoregressive: it generates the entire RVQ token grid in a fixed number of parallel refinement steps, with zero-shot voice cloning and speech editing arising natively as infilling. Luna-TTS Realtime, derived by continual training, is autoregressive over blocks of 32 codec frames (1.28s) while denoising each block in parallel; it supports KV-cached blockwise generation and incremental audio delivery, achieving an end-to-end RTF of 0.0240 and 41.6 ms local first-block latency under the warmed serving protocol. An annealed fine-tuning stage adds explicit control over emotion and non-verbal vocalizations (NVVs), and a reinforcement-learning stage applies GRPO with policy ratios computed over the realized denoising trajectory. On Seed-TTS-Eval, Luna-TTS achieves the best results on all four metrics among compared open-source and commercial systems (0.73 CER / 79.7 SIM on test-zh, 1.49 WER / 76.8 SIM on test-en); on the harder in-the-wild CV3-Eval, it posts the lowest Mandarin and English error rates in our comparison. Against leading commercial systems, it achieves the best results on most objective, model-based, and human-rated metrics for NVV and emotion control.

发表机构

  • VUI Labs Research(VUI实验室研究院)

机构由 AI 辅助整理,请以论文原文为准。

↑