arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

Qwen-音乐技术报告

Qwen-Music Technical Report

Jin Xu, Kangdi Wang, Ruibin Yuan, Shun Lei, Xiong Wang, Xize Cheng, Xueyao Zhang, Yang Zhang, Yiheng Chen, Yongqi Wang, Yue Wang, Zhifang Guo, Zihan Liu, Zijian Lin, Dake Guo, Hangrui Hu, Lei Xie, Linhan Ma, Wei Xue, Wenxiang Guo, Xinfa Zhu, Xipin Wei, Yangze Li, Yuanjun Lv, Yuxuan Wang, Yunfei Chu, Zhiyong Wu

arXiv 2607.11699首次发表:更新:

发表机构

Qwen Team(通义团队)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

介绍Qwen-音乐,它支持文本到音乐生成和翻唱歌曲生成两个核心任务。集成三个核心组件,通过特定机制建模并渲染。经多语言数据训练及多种训练方式,在多个音乐性和音频质量指标上达领先成果,获专业评估人员青睐。

AI 中文摘要

在本报告中,我们介绍了Qwen-音乐,这是一个强大的音乐生成模型,能够生成带有完整人声演唱的高度音乐性和高保真歌曲。Qwen-音乐支持两个核心任务:文本到音乐生成,即根据文本描述、歌词和音乐属性创作全新歌曲;翻唱歌曲生成,即重新诠释具有不同风格和人声特征的现有歌曲。在架构上,Qwen-音乐集成了三个核心组件:Qwen-音乐分词器、Qwen-音乐语言模型和Qwen-音乐渲染器。Qwen-音乐分词器将音频压缩成25赫兹的单码本音乐语义令牌流,为语言模型预测保留语义和旋律信息。基于这些令牌,Qwen-音乐语言模型执行自回归音乐语义建模,关键创新是基于旋律令牌的思维链机制,在整首歌曲生成前规划旋律,提高创造力、音乐性、结构连贯性和基于参考音频的旋律克隆。为克服离散语义令牌的保真度限制,Qwen-音乐渲染器执行生成式立体声渲染,丰富声学细节并产生高保真立体声波形。最后,我们在超过500万小时的涵盖数百种语言的多语言音乐数据上训练Qwen-音乐语言模型。我们首先应用质量感知预训练课程,然后使用渐进式后训练,包括监督初始化、离线DPO和在线GSPO,以进一步提高音乐性和指令跟随能力。在600个中文和英文提示下,Qwen-音乐在16个客观音乐性和音频质量指标中的13个上取得了领先成果。专业评估人员也比领先的专有系统更喜欢Qwen-音乐。对于翻唱歌曲生成,Qwen-音乐比领先的专有系统更准确地保留参考旋律。

英文摘要

We introduce Qwen-Music, a music generation model that produces high-fidelity songs with complete vocals. It supports text-to-music generation from descriptions, lyrics, and musical attributes, and cover song generation with different styles and vocal characteristics. Qwen-Music comprises three components: Qwen-Music-Tokenizer, Qwen-Music-LLM, and Qwen-Music-Render. The tokenizer compresses audio into a 25 Hz single-codebook stream of Music Semantic Tokens that preserve semantic and melodic information. The LLM performs autoregressive modeling with a melody-token-based chain-of-thought (Melody-CoT) mechanism that plans melodies before full-song generation, improving musicality, structural coherence, and reference-melody preservation. The renderer enriches discrete semantic tokens with acoustic details to produce high-fidelity stereo waveforms. We train the LLM using a quality-aware pre-training curriculum followed by progressive post-training with supervised initialization, offline DPO, and online GSPO to improve musicality and instruction following. On 600 evaluation inputs, Qwen-Music achieves state-of-the-art results in 13 of 16 objective musicality and audio-quality metrics. Professional evaluators also prefer Qwen-Music over leading proprietary systems. For cover generation, Qwen-Music preserves reference melodies more accurately than Suno V5.5, Suno V5, and MiniMax Cover on the AI-generated reference set, and outperforms MiniMax Cover on most metrics on the real-world popular-song reference set.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑