arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

哪些约束缺失了?询问验证器:用于约束跟随音乐生成的梯度奖励

Do Language Models Need Music Supervision? Verifiable Rewards for Multi-Constraint Symbolic Music Generation

Haoyue Liu, Xiaoying Tang

arXiv 2609.23665首次发表:更新:

发表机构

The Chinese University of Hong Kong, Shenzhen; Shenzhen Future Network of Intelligence Institute (FNii-Shenzhen)(香港中文大学(深圳); 深圳市未来智联网络研究院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对约束跟随音乐生成,提出MusicRLVR方法,利用验证器提供分级奖励,在MusicConstraintBench上显著提升模型性能,并泛化至未见组合。

AI 中文摘要

约束跟随音乐生成要求乐谱同时满足多个用户指定的属性,每个属性都可以通过程序化方式检查(如调性、拍号、长度、音域、尾音、节奏、律动和结构),但现有基准测试并未单独评估这一能力。我们构建了MusicConstraintBench,包含八个约束族共2180个条目,当前模型在组合几个约束后即表现失败。自然的解决方案是使用这些验证器作为奖励进行强化学习,然而我们观察到,仅在所有属性都满足时才给予奖励,会导致大多数训练组缺乏学习信号:在前50次更新中,0.550的生成组得分完全相同且未获得梯度,尽管失败的得分通常仅缺失一个所要求的属性。在联合准则下,同一提示的生成结果往往共同失败,因此二元奖励无法区分接近正确的乐谱与格式错误的乐谱。为此,我们引入了MusicRLVR,它在拒绝格式错误输出的硬验证门之后,为每个属性提供分级奖励,并附加联合满足奖励,无需人工标注、学习奖励模型或音乐领域微调。在MusicConstraintBench上,MusicRLVR将Qwen3-4B-Instruct在混合约束下的得分从0.160提升至0.807,并超越了包括Llama-3.1-70B(0.380)在内的所有零样本基线。它还能泛化到训练中未见过的属性组合及超出范围的参数值,表明可验证奖励无需预设目标输出。

英文摘要

Language models now generate symbolic music from text, and research has focused on musicality. However, many applications require a score that meets explicit constraints, which models struggle to satisfy jointly: on MusicConstraintBench, our benchmark of 2,180 items over eight families of programmatically verifiable constraints, Llama-3.1-70B satisfies 0.630 of single-constraint items but only 0.044 of four-constraint ones. As a remedy, we introduce MusicRLVR, which trains a language model with group relative policy optimisation (GRPO) on verifier rewards alone, needing no human annotation, reward model or music-domain supervised fine-tuning. MusicRLVR incorporates (1) a hard validation gate that rejects malformed scores, (2) graded per-family credit that, unlike a binary reward, separates partially correct outputs, and (3) an all-satisfied bonus for meeting every constraint at once. Extensive experiments show that, in under four hours of training, MusicRLVR raises Qwen3-4B-Instruct-2507 from 0.160 to 0.797 on mixed constraints, outperforming Llama-3.1-70B, and generalises to unseen property combinations, out-of-range parameters and more constraints than any training prompt. The recipe transfers to Qwen3-8B, and neither trained model loses significant accuracy on general benchmarks.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑