AI 中文总结
StepAudio 3 Gen 是一个基于离散自回归和 RVQ 标记的通用音频生成模型,通过渐进预训练、RVQ 适配器等设计,在 TTS 和声音设计上达到最先进性能,并支持多种音频类型生成。
AI 中文摘要
我们推出了 StepAudio 3 Gen,一个通用音频生成模型,支持零样本文本到语音(TTS)、声音设计、人声生成、音效、音乐、氛围语音以及多种音频类型的混合,所有这些都在一个统一框架内实现。其核心是,StepAudio 3 Gen 是一个离散自回归生成器,直接在残差向量量化(RVQ)标记上对音频进行建模,这与近期通用音频模型中普遍采用的基于扩散 Transformer 的连续生成范式不同。其 StepAudio Tokenizer 在共享的 $16 \ imes 2048$ 残差码空间中,以 12.5 Hz 的频率表示通用音频,联合量化语义和波形级声学特征,使得每个码层都保留两种类型的信息。在生成过程中,主干网络沿时间轴使用自回归建模预测第一个码本,而一个轻量级因果 Transformer 沿码本轴完成其余十五个码本。我们的研究进一步确定了三个关键设计原则:(1)干扰感知的渐进预训练,用于获取音频能力,同时保留大语言模型的文本能力;(2)RVQ 适配器,用于有效整合多码本声学表示;(3)在通用音频领域的共享表示上进行离散自回归建模。通过渐进预训练、多任务指令训练和监督微调,StepAudio 3 Gen 在 TTS 和声音设计方面均达到了最先进的性能,同时在语音、人声、音效和音乐方面保持了强大的生成能力。音频样本可在以下网址获取:此 https URL。
英文摘要
We introduce StepAudio 3 Gen, a general-purpose audio generation model that supports zero-shot text-to-speech (TTS), voice design, vocal generation, sound effects, music, vibe speech, and mixtures of multiple audio types within a unified framework. At its core, StepAudio 3 Gen is a discrete autoregressive generator that models audio directly over residual vector quantization (RVQ) tokens, departing from the diffusion Transformer-based continuous generation paradigm prevalent in recent general audio models. Its StepAudio Tokenizer represents general audio at 12.5 Hz in a shared $16 \times 2048$ residual code space, jointly quantizing semantic and waveform-level acoustic features so that each code layer preserves both types of information. For generation, the backbone predicts the first codebook along the time axis using autoregressive modeling, while a lightweight causal Transformer completes the remaining fifteen codebooks along the codebook axis. Our study further identifies three key design principles: (1) interference-aware progressive pretraining for acquiring audio capabilities while preserving the textual abilities of the large language model, (2) RVQ Adaptor for effectively incorporating multi-codebook acoustic representations, and (3) discrete autoregressive modeling over a shared representation across general audio domains. With progressive pretraining, multi-task instruction training, and supervised fine-tuning, StepAudio 3 Gen achieves state-of-the-art performance on both TTS and voice design, while retaining strong generation capabilities across speech, vocals, sound effects, and music. Audio samples are available at https://stepaudiollm.github.io/step-audio-3-gen/.