AI 中文总结
Diff-Symbo借助LDM与自回归方法,结合含19345个文本模板的数据集,实现了文本控制的长时长高质量符号音乐生成,在多项指标上优于基线模型。
AI 中文摘要
文本控制的符号音乐生成因在音乐创作中具备通用、灵活且直接的方法,近来受到研究关注。然而,以往方法生成的符号音乐往往在质量、多样性、可控性和时长方面存在不足。本文提出Diff-Symbo,一种采用潜在扩散模型(LDM)生成高质量、多样化且长时长符号音乐的创新方法。为解决文本-符号音乐数据集匮乏的问题,我们借助大语言模型开发了包含19345个文本模板的综合数据集。此外,我们设计了音乐信息编码器,以降低训练开销同时提取更有效的控制表征。给定文本描述,所提方法利用LDM提升音乐生成的质量与多样性,还通过自回归方法改善音乐生成的时长与创作一致性。实验结果显示,与GPT-4、MuseCoco和多轨音乐变换器(MMT)等基线模型相比,Diff-Symbo在文本可控性、时长和生成音乐质量方面取得显著提升。作为该领域的先驱模型之一,Diff-Symbo为基于LDM的可控且高质量符号音乐创作铺平了道路,为音乐爱好者与从业者提供了宝贵贡献。
英文摘要
Text-controlled symbolic music generation has recently gained research attention due to its versatile, flexible and straightforward approach to music composition. However, previous approaches tend to generate symbolic music with compromising quality, diversity, controllability and limited duration. In this paper, we present Diff-Symbo, an innovative method that uses latent diffusion model (LDM) to generate high-quality, diverse and long-duration symbolic music. To address the lack of text-symbolic music dataset, we develop a comprehensive dataset with 19,345 text templates by employing large language model. Furthermore, we design a music information encoder to reduce the training overhead while extracting more effective control representations. Given textual descriptions, our proposed method leverages LDM to improve the quality and diversity of music generation. Our method also improves the duration and the compositional consistency of music generation through an autoregressive approach. Experimental results show significant improvements of Diff-Symbo in text controllability, duration, and the quality of generated music compared to the baseline models such as GPT-4, MuseCoco and Multitrack Music Transformer (MMT). As one of the pioneer models in this field, Diff-Symbo paves the way towards controllable and high-quality symbolic music composition based on LDM, offering valuable contributions to both music amateurs and practitioners.