arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.00011cs.CLcs.AI

DLLM-TTS:用于文本到语音合成的块离散扩散语言模型

DLLM-TTS: Block Discrete Diffusion Language Model for Text-to-Speech Synthesis

Wasim Madha, Nityanand Mathur, Hamees Sayed, Apoorv Singh, Sameer Khurana, Akshat Mandloi, Sudarshan Kamath

首次发表
浏览论文内容

中文总结 AI 辅助

本文提出 DLLM-TTS 框架,将 TTS 建模为基于 X-Codec2 的条件块离散扩散,通过块内掩码扩散与并行预测实现高效语音生成,在 Seed-TTS-eval 基准上取得良好性能,兼具实用性与数据效率。

中文摘要 AI 辅助

当前文本到语音(TTS)系统面临权衡:自回归编解码器语言模型生成的语音清晰度极高,但需要大规模模型和训练数据,且需逐 token 解码;而非自回归方法虽提升了速度,却以语言准确性为代价。本文提出 DLLM-TTS,这一框架将 TTS 建模为基于 X-Codec2 神经音频编解码器 token 的条件块离散扩散过程。该模型将序列分解为块,在块内应用掩码扩散,同时逐块处理,学习局部声学连贯性与全局文本-语音对齐。推理时,块内并行 token 预测实现高效生成,实时因子(RTF)为 0.15。在 Seed-TTS-eval 基准上,采用 20K 小时数据训练的 0.6B 参数模型取得了有竞争力的性能,表明块离散扩散语言模型可实现实用且数据高效的语音合成,支持并行生成。

英文摘要

Current text-to-speech systems face a trade-off: autoregres- sive codec language models produce highly intelligible speech but require large-scale models and training data and decode tokens sequentially, while non-autoregressive approaches im- prove speed at the cost of linguistic accuracy. We present DLLM-TTS, a framework that formulates TTS as conditional block discrete diffusion over X-Codec2 neural audio codec to- kens. The model decomposes sequences into blocks and applies masked diffusion within each block while processing blocks se- quentially, learning both local acoustic coherence and global text-speech alignment. During inference, parallel token pre- diction within blocks enables efficient generation with a real- time factor (RTF) of 0.15. A 0.6B-parameter model trained on 20K hours achieves competitive performance on the Seed- TTS-eval benchmark, demonstrating that block discrete diffu- sion language models enable practical and data-efficient speech synthesis with parallel generation.

↑