arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

MusicLayout:用于可控文本生成音乐的显式结构规划

MusicLayout: Explicit Structural Planning for Controllable Text-to-Music Generation

Shuyu Li, Kejun Zhang, Jiahe Lei, Shulei Ji, Zihao Wang, Jiaxing Yu, Wanying Wu, Lei Wang

arXiv 2608.09035首次发表:更新:

发表机构

College of Artificial Intelligence, Zhejiang University; College of Computer Science and Technology, Zhejiang University; Innovation Center of Yangtze River Delta, Zhejiang University; The Chinese University of Hong Kong; Shandong University; Chu Kochen Honors College, Zhejiang University; Ant Group(浙江大学人工智能学院; 浙江大学计算机科学与技术学院; 浙江大学长三角创新中心; 香港中文大学; 山东大学; 浙江大学竺可桢学院; 蚂蚁集团)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对文本生成音乐结构隐含难控的问题,提出MusicLayout显式中间表示,将其集成到统一自回归框架实现布局级控制,实验证实该方法可改进音乐长程结构组织并支持布局级控制。

AI 中文摘要

文本生成音乐领域发展迅速,但当前系统仍主要依赖全局文本提示,导致生成音乐的结构组织隐含,在音频生成前难以检查、控制或修改。为解决该问题,我们提出MusicLayout,这是一种用于控制文本生成音乐中音乐结构的显式中间表示。MusicLayout将音乐作品描述为段落、织体、重复、变体和乐器级编排的时间对齐布局,作为文本意图与生成音乐之间的可解释规划层。我们将MusicLayout集成到基于统一自回归公式构建的文本生成音乐框架中,模型先生成MusicLayout表示,随后在单个序列中基于该表示预测音频标记。生成的MusicLayout可在音频生成前检查和修改,提供布局级结构控制机制。我们通过布局条件生成、布局操控实验和匹配数据 ablation 对MusicLayout进行评估,结果表明显式布局规划可改进长程结构组织并支持布局级控制。

英文摘要

Text-to-music generation has advanced rapidly, but current systems still rely primarily on global text prompts, leaving the structural organization of generated music implicit and difficult to inspect, control, or revise before audio generation. To address this issue, we introduce MusicLayout, an explicit intermediate representation for controlling musical structure in text-to-music generation. MusicLayout describes a musical piece as a time-aligned layout of sections, textures, repetitions, variations, and instrument-level arrangements, serving as an interpretable planning layer between textual intent and the generated music. We integrate MusicLayout into a text-to-music framework built on a unified autoregressive formulation, where the model first generates a MusicLayout representation and subsequently predicts audio tokens conditioned on this representation within a single sequence. The resulting MusicLayout can be inspected and modified prior to audio generation, providing a mechanism for layout-level structural control. We evaluate MusicLayout through layout-conditioned generation, layout manipulation experiments, and matched-data ablations, providing evidence that explicit layout planning can improve long-range structural organization and support layout-level control. We have released the implementation as open source on GitHub at https://github.com/XaryLee/MusicLayout.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑