发表机构
Qwen Team(通义团队)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
介绍Qwen-音乐,它支持文本到音乐生成和翻唱歌曲生成两个核心任务。集成三个核心组件,通过特定机制建模并渲染。经多语言数据训练及多种训练方式,在多个音乐性和音频质量指标上达领先成果,获专业评估人员青睐。
AI 中文摘要
在本报告中,我们介绍了Qwen-音乐,这是一个强大的音乐生成模型,能够生成带有完整人声演唱的高度音乐性和高保真歌曲。Qwen-音乐支持两个核心任务:文本到音乐生成,即根据文本描述、歌词和音乐属性创作全新歌曲;翻唱歌曲生成,即重新诠释具有不同风格和人声特征的现有歌曲。在架构上,Qwen-音乐集成了三个核心组件:Qwen-音乐分词器、Qwen-音乐语言模型和Qwen-音乐渲染器。Qwen-音乐分词器将音频压缩成25赫兹的单码本音乐语义令牌流,为语言模型预测保留语义和旋律信息。基于这些令牌,Qwen-音乐语言模型执行自回归音乐语义建模,关键创新是基于旋律令牌的思维链机制,在整首歌曲生成前规划旋律,提高创造力、音乐性、结构连贯性和基于参考音频的旋律克隆。为克服离散语义令牌的保真度限制,Qwen-音乐渲染器执行生成式立体声渲染,丰富声学细节并产生高保真立体声波形。最后,我们在超过500万小时的涵盖数百种语言的多语言音乐数据上训练Qwen-音乐语言模型。我们首先应用质量感知预训练课程,然后使用渐进式后训练,包括监督初始化、离线DPO和在线GSPO,以进一步提高音乐性和指令跟随能力。在600个中文和英文提示下,Qwen-音乐在16个客观音乐性和音频质量指标中的13个上取得了领先成果。专业评估人员也比领先的专有系统更喜欢Qwen-音乐。对于翻唱歌曲生成,Qwen-音乐比领先的专有系统更准确地保留参考旋律。
英文摘要
We introduce Qwen-Music, a music generation model that produces high-fidelity songs with complete vocals. It supports text-to-music generation from descriptions, lyrics, and musical attributes, and cover song generation with different styles and vocal characteristics. Qwen-Music comprises three components: Qwen-Music-Tokenizer, Qwen-Music-LLM, and Qwen-Music-Render. The tokenizer compresses audio into a 25 Hz single-codebook stream of Music Semantic Tokens that preserve semantic and melodic information. The LLM performs autoregressive modeling with a melody-token-based chain-of-thought (Melody-CoT) mechanism that plans melodies before full-song generation, improving musicality, structural coherence, and reference-melody preservation. The renderer enriches discrete semantic tokens with acoustic details to produce high-fidelity stereo waveforms. We train the LLM using a quality-aware pre-training curriculum followed by progressive post-training with supervised initialization, offline DPO, and online GSPO to improve musicality and instruction following. On 600 evaluation inputs, Qwen-Music achieves state-of-the-art results in 13 of 16 objective musicality and audio-quality metrics. Professional evaluators also prefer Qwen-Music over leading proprietary systems. For cover generation, Qwen-Music preserves reference melodies more accurately than Suno V5.5, Suno V5, and MiniMax Cover on the AI-generated reference set, and outperforms MiniMax Cover on most metrics on the real-world popular-song reference set.