arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

GEPARD — 面向实时对话的生成式、韵律感知自回归文本转语音模型

GEPARD - Generative, Prosody-aware, Autoregressive text-to-speech model for Realtime Dialogue

Denis Pavlov, Ulanbek Abdurazakov, Nursultan Bakashov

arXiv 2609.04222首次发表:更新:

AI 中文总结

本研究提出面向实时对话的流式TTS模型GEPARD,以vLLM为骨干,通过将辅助机制移至预填充或蒸馏至权重实现高效推理,单流实时因子约0.067,256并发流总加速比约204倍,还解决了短寄存器失效等问题。

AI 中文摘要

我们提出了GEPARD(Generative, Prosody-aware, Autoregressive text-to-speech model for Realtime Dialogue),这是一种用于实时口语对话的流式文本转语音模型。GEPARD以大语言模型(LLM)作为骨干进行自回归语音生成——文本与音频嵌入在单一的仅解码器模型中联合训练,并通过基于FSQ的神经编解码器将其解码为波形,随文本的到来逐块流式生成音频。我们的核心目标是构建一种由标准LLM引擎(vLLM)提供服务的TTS架构,且无需修改其计算内核。这定义了总体设计原则:骨干为标准的全注意力Transformer,而所有非平凡的辅助机制——零样本语音克隆、文本增强以及无分类器引导——均被移出自回归解码循环,移至预填充阶段,或直接蒸馏至模型权重中。在流式端到端推理中,单条流的实时因子约为0.067(比实时速度快约15倍);在单台服务器级GPU上,256条并发流下,系统的总加速比约达204倍。我们详述了:(1)适配vLLM原生服务的系统级解决方案;(2)自回归语音解码器的“短寄存器”(1-2个词)失效模式,以及诊断探针与缓解方法;(3)通过直接偏好优化(DPO)将文本上的两阶段无分类器引导蒸馏为单阶段权重。

英文摘要

We present GEPARD (Generative, Prosody-aware, Autoregressive text-to-speech model for Realtime Dialogue), a streaming text-to-speech model for real-time spoken dialogue. GEPARD generates speech autoregressively with an LLM backbone - text and audio embeddings are trained together in a single decoder-only model - and decodes it to a waveform with an FSQ-based neural codec, streaming audio chunk-by-chunk as text arrives. Our central goal is a TTS architecture served by a standard LLM engine (vLLM) without modifying its compute kernels. This defines the overarching design principle: the backbone is a standard full-attention transformer, while all non-trivial auxiliary mechanisms - zero-shot voice cloning, text augmentation, and classifier-free guidance - are moved out of the autoregressive decode loop into prefill, or distilled directly into the weights. On streaming end-to-end inference, a single stream reaches a Real-Time Factor of about 0.067 (roughly 15x faster than real-time); under 256 concurrent streams the system reaches an aggregate speedup of about 204x on a single server-class GPU. We detail: (1) system-level solutions for vLLM-native serving; (2) the "short register" (1-2 word) failure mode of autoregressive speech decoders, with diagnostic probes and a mitigation; and (3) distillation of two-pass classifier-free guidance over text into single-pass weights via Direct Preference Optimization (DPO).

CommentsTechnical Report. 37 pages, 11 figures, Demo samples, code, and open weights: https://huggingface.co/nineninesix/gepard-1.0 ; https://github.com/nineninesix-ai/gepard-inference . Affiliation: Nineninesix, Inc

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑