arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.16034eess.AScs.SD

StepAudio 3 音乐技术报告

StepAudio 3 Music Technical Report

  • StepFun ACE
  • The Chinese University of Hong Kong(香港中文大学)
  • University of California San Diego(加利福尼亚大学圣迭戈分校)

机构由 AI 辅助整理,请以论文原文为准。

Chengli Feng, Zhiyue Wu, Jiahao Song, Zheqi Dai, Boyang Wang, Ruibin Yuan, Junming Gong, Wenxiao Zhao, Jing Guo, Gang Yu, Xiangyu Zhang, Xuerui Yang, Chao Yan

AI总结:

StepAudio 3 Music 通过离散-连续设计、ABC 记谱法显式规划和 DPO 强化学习,实现长达 5 分 30 秒的高质量长格式音乐生成,在多项基准上取得领先。

AI中文摘要:

我们推出了 StepAudio 3 Music,一个大规模、长格式音乐生成模型,支持显式音乐规划和开放域文本控制生成。StepAudio 音乐分词器将音频表示为来自 65536 条目单码本的 50-Hz 流,利用语义信息的自监督和多任务训练来保留音乐结构和重建相关信息。一个流匹配扩散 Transformer(DiT)预测连续的 StepAudio VAE 潜变量,我们的 VAE 解码器将其转换为 48-kHz 音频。这种离散-连续设计通过单码本 VQ、语义和声学 RVQ 以及不同 DiT 配置的比较来指导。对于显式规划,一个专家混合自回归模型使用 ABC 记谱法在预测音乐令牌之前生成中间编排计划(ABC-CoT),使和声、节奏和旋律结构成为生成上下文的一部分。渐进式训练课程和监督微调支持歌曲和器乐生成、从干人声生成伴奏,以及长达 5 分 30 秒的翻唱歌曲合成。通过直接偏好优化(DPO)的强化学习,最终模型在评估系统中获得了最高的 AudioBox 内容享受度、内容有用性和制作质量分数,以及最高的 MuQ-MuLan 相似度,SongBench 结果具有竞争力。在初步的人工分析音乐竞技场人声排行榜上,它获得了 1105 的质量 Elo,仅次于 Suno V5.5 和 Mureka,领先于 Suno V5、MiniMax 模型和其他系统。音频演示可在该 https URL 获取。

英文摘要:

We introduce StepAudio 3 Music, a large-scale, long-form music generation model that supports explicit musical planning and open-domain text-controlled generation. The StepAudio Music Tokenizer represents audio as a 50-Hz stream from a 65536-entry single codebook, using semantically informed self-supervised and multi-task training to preserve musical structure and reconstruction-relevant information. A flow-matching diffusion Transformer (DiT) predicts continuous StepAudio VAE latents, which our VAE decoder converts into 48-kHz audio. This discrete-continuous design is guided by comparisons of single-codebook VQ, Semantic and Acoustic RVQ, and different DiT configurations. For explicit planning, a Mixture-of-Experts autoregressive model uses ABC notation to produce an intermediate arrangement plan (ABC-CoT) before predicting music tokens, making harmony, rhythm, and melodic structure part of the generation context. A progressive training curriculum and supervised fine-tuning support song and instrumental generation, accompaniment generation from dry vocals, and cover-song synthesis for up to 5 minutes and 30 seconds. With reinforcement learning via direct preference optimization (DPO), the final model achieves the highest AudioBox Content Enjoyment, Content Usefulness, and Production Quality scores and the highest MuQ-MuLan similarity among the evaluated systems, with competitive SongBench results. On the preliminary Artificial Analysis Music Arena Vocals leaderboard, it obtains a Quality Elo of 1105, behind only Suno V5.5 and Mureka and ahead of Suno V5, MiniMax models, and other systems. Audio demonstrations are available at https://stepaudiollm.github.io/step-audio-3-music.

补充信息

相关深度报道

↑