arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

P-MUSE:用于统一器乐合成与编辑的提示-MIDI可选模型

P-MUSE: Prompt-MIDI-Optional Model for Unified Instrumental Music Synthesis and Editing

Chong Jing, Junan Zhang, Jing Yang, Yulun Wu, Fan Fan, Zhizheng Wu

arXiv 2608.01920首次发表:更新:

AI 中文总结

P-MUSE是统一器乐合成与编辑的MIDI-to-Music框架,通过多阶段课程学习整合两种生成范式,提出相位感知引导策略与尾端丢弃法,并建立含四类乐器的综合基准。

AI 中文摘要

MIDI-to-Music系统可将目标MIDI序列的旋律与节奏渲染为音乐片段,同时从提示录音中克隆乐器音色。现有系统通常采用两种不同范式之一:仅基于提示音频的条件生成,适用于缺少对齐提示MIDI的场景;以及基于配对提示音频与MIDI的上下文学习,利用跨模态对齐实现对MIDI跟随与音色相似度的更强控制。我们引入P-MUSE,这一器乐MIDI-to-Music框架通过支持提示-MIDI可选输入的多阶段课程学习,统一了上述两种范式。P-MUSE还通过共享的中间填充公式统一了音乐生成与局部编辑。基于理论分析与实证研究,我们为转录到音频系统提出了相位感知的无分类器引导调度原则,以及尾端丢弃(Tail-Drop)策略。最后,为推动该领域研究,我们建立了首个综合基准,涵盖多种提示模式、生成/编辑任务,以及四种代表性乐器:钢琴、吉他、贝斯与鼓。演示可访问此https URL查看。

英文摘要

MIDI-to-Music system renders the melody and rhythm of a target MIDI sequence into musical segment while cloning instrument timbre from a prompt recording. Existing systems typically adopt one of two distinct paradigms: conditional generation with prompt audio alone, which remains applicable when aligned prompt MIDI is unavailable, and In-Context Learning with paired prompt audio and MIDI, which exploits cross-modal alignment for stronger control on MIDI following and timbre similarity. We introduce P-MUSE, an instrumental MIDI-to-Music framework that unifies both paradigms via a multi-stage Curriculum-Learning supporting prompt-MIDI-optional inputs. P-MUSE further unifies music generation and local editing through a shared fill-in-the-middle formulation. Grounded in theoretical analysis and empirical study, we propose a phase-aware classifier-free guidance scheduling principle for Transcription-to-Audio systems, alongside a Tail-Drop strategy. Finally, to advance research in this field, we establish the first comprehensive benchmark, covering various prompt modes, generation/editing tasks, and four representative instruments: piano, guitar, bass, and drums. Demos are available at https://p-muse.github.io/.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑