从一次性生成到增量式音乐创作:使通用指令大语言模型适应持久化符号编辑
From One-Shot Generation to Incremental Music Composition: Adapting a General-Purpose Instruction LLM for Persistent Symbolic Editing
浏览论文内容
中文总结 AI 辅助
本研究将通用指令大语言模型改造为可复用的符号乐谱编辑操作符,通过操作感知状态转换和LoRA微调,使检查器通过率从29.37%提升至99.37%,验证了增量式作曲的技术可行性。
中文摘要 AI 辅助
大多数音乐生成系统仍主要被构建和评估为完整输出的生产者,而作曲过程通常是通过对共享音乐作品的连续修订来进行的。本文研究了一种通用指令跟随大语言模型的不同用途:不是作为一次性音乐生成器,而是作为可复用的操作符作用于不断演化的符号乐谱。我们将增量式作曲形式化为对持久化ABC记谱法的一系列操作感知状态转换,并明确规定了每个操作可以改变什么以及必须保留什么。交互包括两种作品初始化变体和三种编辑操作——和弦添加、修补和移调。我们通过使用低秩适配(LoRA)对Llama 3.1 8B Instruct进行适配来实例化该形式化方法,训练数据为源自爱尔兰传统音乐的496,038条操作感知对话记录。与未适配模型的比较用于测试学习这种交互契约的可行性,而非声称微调本身的新颖性。在每模型500个对话(共1,750个尝试输出状态)中,检查器通过率从29.37%提升至99.37%,而在通过条件下的合规率从0.7205提升至0.9798。符合参考相对音乐特征分析严格资格的输出从14个增加到1,548个,最长公共子序列分析未显示在指定协议下相对于保留基线的高重叠序列的系统性增加。结果支持了使用通用指令大语言模型进行持久化、操作感知符号编辑的技术可行性。这些结果并未确立更优的音乐质量或人机共创性,这些问题仍有待以音乐家为中心的评估来解决。
英文摘要
Most music-generation systems are still framed and evaluated primarily as producers of complete outputs, whereas composition often proceeds through successive revisions to a shared musical artifact. This paper studies a different use of a general-purpose instruction-following large language model: not as a one-shot music generator, but as a reusable operator over an evolving symbolic score. We formulate incremental composition as a sequence of operation-aware state transitions over persistent ABC notation, with explicit requirements on what each operation may change and what it must preserve. The interaction includes two artifact-initialization variants and three editing operations -- chord addition, inpainting, and transposition. We instantiate the formulation by adapting Llama 3.1 8B Instruct with Low-Rank Adaptation (LoRA) on 496,038 operation-aware dialogue records derived from Irish traditional music. The comparison with the unadapted model is used to test the feasibility of learning this interaction contract, not to claim novelty for fine-tuning itself. Across 500 dialogues per model (1,750 attempted output states), checker admission rises from 29.37% to 99.37%, while compliance conditional on admission rises from 0.7205 to 0.9798. Strict eligibility for reference-relative musical-feature analysis increases from 14 to 1,548 outputs, and Longest Common Subsequence analysis does not show a systematic increase in high-overlap sequences relative to held-out baselines under the specified protocol. The results support the technical feasibility of persistent, operation-aware symbolic editing with a general-purpose instruction LLM. They do not establish superior musical quality or human-AI co-creativity, which remain questions for musician-centered evaluation.
发表机构
- Pontifícia Universidade Católica do Rio de Janeiro (PUC-Rio)(里约热内卢天主教大学)
- Sorbonne Université(索邦大学)
- LIP6(巴黎第六大学计算机科学实验室)
- Universidade Federal do Rio de Janeiro (UFRJ)(里约热内卢联邦大学)
机构由 AI 辅助整理,请以论文原文为准。