发表机构
Dartmouth College(达特茅斯学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究针对音乐后期制作问题,引入RIME框架,利用POEMS工具包生成配对数据,评估多模态语言模型,展示其通过监督微调提升智能体性能,是迈向迭代音乐智能体、变革音乐制作的早期步骤。
AI 中文摘要
几乎每一首你听过的录制音乐在到达你耳中之前都经过了修改;商业发行的音乐很少完全由音乐家脑海中直接诞生。尽管音乐生成模型有望实现一次性输出,但这种细粒度的迭代改进工作流程是一个互补问题,且模型大多难以企及。音乐家也面临差距:他们能表达想听的,但并非所有人都具备使用录音室制作工具来实现这些直觉所需的复杂操作能力。我们将此任务形式化为智能后期制作,即针对歌曲的各个方面进行改进并组合成最终曲目。我们认为瓶颈在于数据:现有语料库未反映现实后期制作链与音乐家和工程师实际使用词汇的对应关系。我们提出存在一种用于修改录制音乐的密集、一致且可学习的语言。我们引入基于规则的音乐编辑指令(RIME)框架,它能从任何基于规范方法、设计模式和实际制作工作流程约束的基线音乐数据集中生成逼真的配对编辑指令数据。RIME利用POEMS这一新工具包,它结合了音轨分离、混音和常见录音室效果供多模态智能体使用。我们使用RIME和POEMS生成3000对编辑指令和真实音频,并以此数据评估现有多模态语言模型在此任务中的智能体表现,显示当前模型在后期制作能力方面存在持续挑战。我们还展示了RIME通过监督微调提高后期制作智能体性能的能力。我们将RIME视为迈向迭代音乐智能体的早期步骤,这种协作系统可能会像交互式编码智能体重塑软件工程一样改变音乐制作。
英文摘要
Almost every piece of recorded music you have ever heard was modified before it reached you; commercial releases rarely spring fully formed from a musician's mind. Despite the promise of music generation models for one-shot output, such fine-grained iterative refinement workflows are a complementary problem and largely out of their reach. There is also a gap for musicians: while they can express what they want to hear, not all have the facility with studio production tools to implement the complex set of actions needed to realize these intuitions. We formalize this task as agentic post-production, wherein individual aspects of a track are targeted, refined, and combined into a final version. We argue the bottleneck is data: existing corpora do not reflect how realistic post-production chains map onto the vocabulary musicians and engineers actually use. We observe that there is a language for modifying recorded music that is dense, consistent, and learnable. To leverage this, we introduce the Rule-based Instructions for Music Editing (RIME) framework, which generates realistic paired edit-instruction data from any baseline music dataset grounded in canonical methods, design patterns, and constraints derived from real production workflows. RIME leverages POEMS, a toolkit that combines stem separation, mixing, and common studio effects for use by multimodal agents. We use POEMS and RIME to generate 15,000 pairs of edit instructions and ground-truth audio, then use this data to evaluate existing multimodal LLMs as agents on this task, revealing persistent limitations in current models. We also demonstrate RIME's ability to improve post-production agent performance via supervised fine-tuning on synthetic data. We see RIME as a step towards iterative musical agents, collaborative systems that could transform music production much as interactive coding agents have reshaped software engineering.