arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.33757eess.AScs.LGcs.SD

YuE2:在前沿质量上统一符号音乐与音频音乐生成

YuE2: Unifying Symbolic and Audio Music Generation at Frontier Quality

  • HKUST(香港科技大学)
  • M-A-P(M-A-P(多模态艺术投影))
  • New York University(纽约大学)
  • Stanford University(斯坦福大学)
  • MBZUAI(穆罕默德·本·扎耶德人工智能大学)
  • ACE Studio(ACE工作室)
  • HKGAI(香港生成式人工智能研究院)

机构由 AI 辅助整理,请以论文原文为准。

Ruibin Yuan, Jiahao Pan, Junyan Jiang, Zhiyue Wu, Ziya Zhou, Jiankai Sun, Yizhi Li, Ge Zhang, Yicheng Gu, Zeyue Tian, Junyu Dai, Hanfeng Lin, Kai Li, Shangda Wu… 展开作者

Ruibin Yuan, Jiahao Pan, Junyan Jiang, Zhiyue Wu, Ziya Zhou, Jiankai Sun, Yizhi Li, Ge Zhang, Yicheng Gu, Zeyue Tian, Junyu Dai, Hanfeng Lin, Kai Li, Shangda Wu, Xuanjie Liu, Jiaming Wang, Zihan Liu, Yue Wang, Yinghao Ma, Hanzhi Yin, Kangrui Chen, Xinyue Zhang, Ziyang Ma, Mengqi Liao, Hejia Zhao, Guowei Huang, Chao Yan, Lei Ke, Jianwei Yu, Bei Liu, Joe Guo, Liumeng Xue, Gus Xia, Wei Xue, Yike Guo

AI总结:

YuE2通过符号规划统一符号与音频音乐生成,采用AR-NAR混合Transformer,实现前沿质量,并在多项基准上超越现有系统,支持编辑和零样本翻唱。

AI中文摘要:

符号模型能够明确地生成旋律、和声、节奏和曲式,但通常止步于完成录音之前;音频模型能够生成完整歌曲,却将作曲过程隐式化。我们提出YuE2,通过符号规划在前沿质量上统一符号音乐与音频音乐生成。一个单一的AR-NAR混合Transformer(MoT)首先写出可读的乐谱,指定旋律与和声,将其扩展为语义音乐令牌,并实现为完整歌曲的音频。在使用同一检查点的对比中,专家更偏好符号规划带来的整体质量和音乐性,总体偏好为49.3%,而无规划时为34.6%。专家还更偏好统一模型,而非单独的语音模型和扩散Transformer。在WildSongBench上,YuE2在SongBench全局平均上得分为6.73,超过了所有评估的公开基线。从八个候选中选择(best-of-8),YuE2达到6.96,是所有评估系统中最高的观测均值。专家聆听进一步确立了其与专有歌曲生成器的竞争力,best-of-8优于Suno v4.5,并与Suno v5相比获得几乎平衡的偏好。为了从没有对齐乐谱的录音中学习这一生成过程,我们引入了MERT2和SheetSage2,以提供语义和符号监督。MERT2在音乐表示学习上创下了新的最先进水平,在15个MARBLE指标中超过了14个之前的最高结果;SheetSage2在我们的主谱转录比较中,在15个基准-指标对中领先12个。同一检查点能遵循乐谱编辑,同时基本保留未编辑的音乐内容,并能生成零样本翻唱,无需特定于翻唱的训练。其可读乐谱还支持智能体音乐编辑,外部语言模型将用户反馈转化为对作曲的修订。

英文摘要:

Symbolic models make melody, harmony, rhythm, and form explicit but typically stop before a finished recording; audio models produce complete songs while leaving composition implicit. We introduce YuE2, which unifies symbolic and audio music generation at frontier quality through symbolic planning. A single AR-NAR Mixture-of-Transformers (MoT) first writes a readable score specifying melody and harmony, expands it into semantic music tokens, and realizes it as full-song audio. In comparisons using the same checkpoint, experts prefer symbolic planning for overall quality and musicality, with 49.3% of overall preferences versus 34.6% without planning. Experts also favor the unified model over a separate language model and diffusion Transformer. On WildSongBench, YuE2 scores 6.73 on SongBench Global Avg, exceeding all evaluated public baselines. Selecting from eight candidates (best-of-8), YuE2 reaches 6.96, the highest observed mean among all evaluated systems. Expert listening further establishes its competitiveness with proprietary song generators, favoring best-of-8 over Suno v4.5 and yielding nearly balanced preferences against Suno v5. To learn this generation process from recordings without aligned scores, we introduce MERT2 and SheetSage2 to supply semantic and symbolic supervision. MERT2 sets a new state of the art in music representation learning, surpassing previous best results on 14 of 15 MARBLE metrics; SheetSage2 leads 12 of 15 benchmark-metric pairs in our lead-sheet transcription comparison. The same checkpoint follows score edits while largely preserving unedited musical content and generates zero-shot covers without cover-specific training. Its readable score also enables agentic music editing, with external language models translating user feedback into revisions of the composition.

补充信息

↑