TontaubeV1:具有分层编解码器建模和有界上下文的流式文本到语音合成
TontaubeV1: Streaming Text-to-Speech with Hierarchical Codec Modeling and Bounded Context
浏览论文内容
中文总结 AI 辅助
TontaubeV1提出分层编解码器与有界上下文流式TTS,在单GPU上实现自然韵律与低延迟,匹配ElevenLabs并超越多个竞品。
中文摘要 AI 辅助
文本到语音系统通常在自然韵律和高效推理之间面临权衡:更高的感知质量通常伴随着更高的计算成本和延迟。我们提出了TontaubeV1,一个在保持自然韵律的同时,能够在单个消费级GPU上进行流式生成的模型。语音由12.5 Hz的分层DualCodec表示编码,该表示将语义流与连续的声学细化分离。我们的设计假设韵律结构在语义流生成时已基本建立,并据此分配容量:一个基于Qwen3-1.7B的Transformer预测该语义流,从而预测话语时长,而三个逐渐变小的基于Qwen3-0.6B的Transformer各自添加一个声学细化。文本按字符而非子词进行分词。在共享位置配对文本和音频标记,支持有界上下文的长文本生成,并将重叠的DualCodec重建映射到VibeVoice声学潜在空间并进行因果解码,从而在DualCodec的非因果解码器下实现流式生成。该模型最多接受一分钟的参考音频进行声音条件化,主要面向英语和德语设计,并支持额外的多语言。四个预测器总计2.9B参数;在单个RTX 5090上,流式路径达到约200毫秒的首音频延迟。在独立的非流式测量中,单输入端到端实时因子(RTF)为0.08,八个并发输入的总RTF为0.02。在我们基于LLM作为评判者的有声书阅读基准上,TontaubeV1在韵律方面与ElevenLabs Flash v2.5相当,并优于Fish Audio S2 Pro、2026年4月的Gradium API和Cartesia Sonic 3。模型权重已在Hugging Face上以Tontaube社区模型许可证1.0发布。
英文摘要
Text-to-speech systems often face a trade-off between natural prosody and efficient inference: higher perceptual quality typically comes at increased computational cost and latency. We present TontaubeV1, a model that preserves natural prosody while enabling streaming from a single consumer GPU. Speech is encoded by the hierarchical DualCodec representation at 12.5 Hz, which separates a semantic stream from successive acoustic refinements. Our design assumes that prosodic structure is largely established when the semantic stream is generated, and allocates capacity accordingly: a Qwen3-1.7B-derived transformer predicts that stream and thereby the utterance duration, while three progressively smaller Qwen3-0.6B-derived transformers each add one acoustic refinement. Text is tokenized per character rather than by subword. Paired text and audio markers at shared positions support long-form generation with bounded context, and overlapping DualCodec reconstructions are mapped into the VibeVoice acoustic latent space and decoded causally, enabling streaming despite DualCodec's noncausal decoder. The model accepts up to one minute of reference audio for voice conditioning and is designed primarily for English and German, with additional multilingual support. The four predictors total 2.9B parameters; on a single RTX 5090 the streaming path reaches approximately 200 ms to first audio. In separate non-streaming measurements, the end-to-end real-time factor (RTF) is 0.08 for one input and the aggregate RTF is 0.02 across eight concurrent inputs. On our LLM-as-a-judge audiobook-reading benchmark, TontaubeV1 matches ElevenLabs Flash v2.5 and outperforms Fish Audio S2 Pro, the April 2026 Gradium API, and Cartesia Sonic 3 on prosody. The model weights are released on Hugging Face under the Tontaube Community Model License 1.0.
发表机构
- Craitech
机构由 AI 辅助整理,请以论文原文为准。