arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

推动全曲生成的前沿:分层自回归规划与流匹配渲染

Pushing the Frontier of Full-Song Generation: Hierarchical Autoregressive Planning Meets Flow-Matching Rendering

Junyu Dai, Xinyue Fan, Weiqin Li, Xiangang Li, Yunjia Li, Bin Ma, Yukun Ma, Chongjia Ni, Yufei Shi, Biao Tian, Haoxu Wang, Menglin Wu, Jianwei Yu, Huaicheng Zhang, Han Zhao, Shengkui Zhao, Haina Zhu

arXiv 2607.20253首次发表:更新:

发表机构

Alibaba(阿里巴巴)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究提出统一歌曲生成框架,支持多项任务,由四个组件构成。通过特定编码、建模、匹配及模块实现歌曲生成,还研究多种后训练策略,实验表明该框架在评估中性能具有竞争力。

AI 中文摘要

在本报告中,我们提出了一个统一的歌曲生成框架,能够从歌词、文本描述和音乐属性中生成高质量的全长音乐。该框架支持三项任务:歌词到歌曲生成、器乐音乐生成和翻唱歌曲生成。系统架构上由语义感知分词器、hybird-LM、FullDiT和两级旋律模块四个主要组件组成。分词器将音频编码为8码本RVQ令牌,hybird-LM进行分层自回归音频令牌建模,FullDiT在连续VAE潜在空间中进行全曲流匹配,旋律模块用于翻唱歌曲生成。最后研究了DPO、GRPO和OPD作为基于奖励的后训练策略,并将基于流的GRPO应用于FullDiT。实验结果表明该框架在评估设置中取得了有竞争力的性能。

英文摘要

In this report, we present a unified song generation framework capable of producing high-quality full-length music from lyrics, text descriptions, and musical attributes. The proposed framework supports three tasks: Lyrics-to-Song Generation, which generates complete songs from text descriptions, lyrics, and musical attributes; Instrumental Music Generation, which creates music without vocals; and Cover Song Generation, which reinterprets existing songs with different styles while preserving their melodic content. Architecturally, our system consists of four main components: a semantic-aware tokenizer, hybird-LM, FullDiT, and a two-level melody module. The tokenizer encodes audio into 8-codebook RVQ tokens for efficient discrete music representation. Based on these tokens, hybird-LM performs hierarchical autoregressive audio-token modeling for full-song generation. To improve audio fidelity, FullDiT performs full-song flow matching in a continuous VAE latent space conditioned on codec tokens, lyrics, and text captions. For cover song generation, the melody module extracts and discretizes melody cues from reference audio to guide generation while preserving the original melodic content. Finally, we investigate DPO, GRPO, and OPD as reward-based post-training strategies for hybird-LM and apply flow-based GRPO to FullDiT to improve musicality and rendering quality. Experimental results on a multilingual automatic benchmark, complemented by the Artificial Analysis Music with Vocals leaderboard, show that the proposed framework achieves competitive performance in the evaluated settings.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑