arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

StepAudio 3 Gen 技术报告

StepAudio 3 Gen Technical Report

Bin Lin, Bo Zhao, Boyang Wang, Boyang Zhang, Boyong Wu, Chao Yan, Chen Geng, Chen Wu, Cheng Yi, Chengli Feng, Chenglin Zhu, DanNi Wan, Daxin Jiang, Dongqing Pang, Fei Tian, Feng Tian, Future Li, Gang Yu, Guanglong Yang, Jia Peng, Jiahao Song, Jiamin Fan, Jiangjie Zhen, Jianzheng Gao, Jun Chen, Li Xie, Lifang Zhang, Lingli Ji, Liying Shi, Lun Cai, Min Xu, Na Wang, Peilin Li, Peng Yang, Pengfei Tan, Qingjian Lin, Ruijie Xiong, Runze Li, Shenghua Hu, Shi Qiu, Siqi Tu, Siyi Zhou, Tianjiao Deng, Wanying Lu, Weiming Niu, Wen Sun, WenWen Qu, Xiangyu Zhang, Xianwei Zhang, XiaoSu Su, Xing Chen, Xinyu Liu, Xuerui Yang, Yang Li, Yang Yang, Yechang Huang, Yibo Zhu, Yifan Zhang, Yiyang Xu, Yu Fu, Yu Luo, Yu Zhou, Yumang Wang, Yunzhou Ju, Yuxiang Yang, Zekai Liu, Zengwei Yao, Zhenwei Mou, Zheqi Dai, Zhiyue Wu, Zichao Zhou

arXiv 2609.12945首次发表:更新:

AI 中文总结

StepAudio 3 Gen 是一个基于离散自回归和 RVQ 标记的通用音频生成模型,通过渐进预训练、RVQ 适配器等设计,在 TTS 和声音设计上达到最先进性能,并支持多种音频类型生成。

AI 中文摘要

我们推出了 StepAudio 3 Gen,一个通用音频生成模型,支持零样本文本到语音(TTS)、声音设计、人声生成、音效、音乐、氛围语音以及多种音频类型的混合,所有这些都在一个统一框架内实现。其核心是,StepAudio 3 Gen 是一个离散自回归生成器,直接在残差向量量化(RVQ)标记上对音频进行建模,这与近期通用音频模型中普遍采用的基于扩散 Transformer 的连续生成范式不同。其 StepAudio Tokenizer 在共享的 $16 \ imes 2048$ 残差码空间中,以 12.5 Hz 的频率表示通用音频,联合量化语义和波形级声学特征,使得每个码层都保留两种类型的信息。在生成过程中,主干网络沿时间轴使用自回归建模预测第一个码本,而一个轻量级因果 Transformer 沿码本轴完成其余十五个码本。我们的研究进一步确定了三个关键设计原则:(1)干扰感知的渐进预训练,用于获取音频能力,同时保留大语言模型的文本能力;(2)RVQ 适配器,用于有效整合多码本声学表示;(3)在通用音频领域的共享表示上进行离散自回归建模。通过渐进预训练、多任务指令训练和监督微调,StepAudio 3 Gen 在 TTS 和声音设计方面均达到了最先进的性能,同时在语音、人声、音效和音乐方面保持了强大的生成能力。音频样本可在以下网址获取:此 https URL。

英文摘要

We introduce StepAudio 3 Gen, a general-purpose audio generation model that supports zero-shot text-to-speech (TTS), voice design, vocal generation, sound effects, music, vibe speech, and mixtures of multiple audio types within a unified framework. At its core, StepAudio 3 Gen is a discrete autoregressive generator that models audio directly over residual vector quantization (RVQ) tokens, departing from the diffusion Transformer-based continuous generation paradigm prevalent in recent general audio models. Its StepAudio Tokenizer represents general audio at 12.5 Hz in a shared $16 \times 2048$ residual code space, jointly quantizing semantic and waveform-level acoustic features so that each code layer preserves both types of information. For generation, the backbone predicts the first codebook along the time axis using autoregressive modeling, while a lightweight causal Transformer completes the remaining fifteen codebooks along the codebook axis. Our study further identifies three key design principles: (1) interference-aware progressive pretraining for acquiring audio capabilities while preserving the textual abilities of the large language model, (2) RVQ Adaptor for effectively incorporating multi-codebook acoustic representations, and (3) discrete autoregressive modeling over a shared representation across general audio domains. With progressive pretraining, multi-task instruction training, and supervised fine-tuning, StepAudio 3 Gen achieves state-of-the-art performance on both TTS and voice design, while retaining strong generation capabilities across speech, vocals, sound effects, and music. Audio samples are available at https://stepaudiollm.github.io/step-audio-3-gen/.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑