发表机构
The University of Osaka; The University of Tokyo(大阪大学; 东京大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究提出一种用于音乐生成的迭代上下文学习框架,发现仅靠排序可在有限条件下指导生成,但可迁移的标准推断受限于基础模型的音乐属性处理能力。
AI 中文摘要
使生成式音乐系统适应个人品味需要学习该听众的偏好标准,听众可对作品进行排序,但潜在标准可能是隐性且难以明确表达的。本研究探究大型语言模型(LLM)是否能仅通过排序来适配符号音乐生成,并构建可迁移的价值标准自然语言描述,以及探究其适配的条件。在本迭代上下文学习框架中,LLM会提出假设、以ABC记谱法生成候选作品、接收排序结果,并定期从历史记录中推断并表述价值标准以指导后续生成。我们通过混合效应模型、消融实验及对未见过的音乐的迁移测试,在480次适配运行中针对16个模拟评分者对该框架进行评估。总体而言,该框架未优于无反馈多样化生成基线,但在两个通过简单采样难以达到目标的价值函数上表现更优。目标相对于LLM无反馈生成倾向的非典型性可预测适配难度;此外,适配过程中的较高价值并不意味着能将该标准识别为通用规则。在未见过的音乐上,获取的描述与历史记录对更多价值函数的生成有改善,其对偏好预测的改善则弱,预测结果仍接近随机水平;部分增益与上下文音乐的声学接近性相关,其余则不然。这些发现表明,仅靠排序可在有限条件下指导生成,而可迁移的标准推断仍受限于基础模型识别、推理及表述音乐属性的能力。
英文摘要
Adapting a generative music system to an individual's taste requires learning what that listener values. Listeners can rank pieces, but their underlying criteria may be tacit and difficult to articulate. We ask whether and under what conditions a large language model (LLM) can adapt symbolic music generation from rankings alone and construct transferable natural-language descriptions of value criteria. In our iterative in-context learning framework, the LLM formulates hypotheses, generates candidate pieces in ABC notation, receives a ranking, and periodically infers and verbalizes value criteria from history to guide later generation. We evaluate the framework against 16 simulated raters in 480 adaptation runs using mixed-effects modeling, an ablation, and transfer tests on unseen music. Overall, the framework did not outperform a feedback-free diverse-generation baseline, but did so for two value functions with targets difficult to reach through simple sampling. How atypical the target was relative to the LLM's feedback-free generation tendencies predicted adaptation difficulty. Moreover, higher value during adaptation did not imply identification of the criterion as a general rule. On unseen music, acquired descriptions and histories improved generation for more value functions than they improved preference prediction, which remained near chance. Some gains were associated with acoustic proximity to music in the context, but others were not. These findings show that rankings alone can guide generation under limited conditions, while transferable criterion inference remains constrained by the foundation model's ability to recognize, reason about, and verbalize musical attributes.