arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2610.03589cs.SD

基于评分标准的文本到音乐生成优化

Rubric-Based Optimization for Text-to-Music Generation

Ping Wang, Guang Yang, Shao-Rong Su, Junkai Wu, Pang Wei Koh, Noah A. Smith

首次发表
浏览论文内容

中文总结 AI 辅助

本研究提出利用音频语言模型的结构化评分标准作为奖励信号,通过DPO和DiffusionNFT优化文本到音乐生成,实验表明其提升整体感知质量,而专门客观奖励更适合精确属性优化。

中文摘要 AI 辅助

后训练阶段的文本到音乐生成需要能够捕捉音乐质量多个方面的奖励信号,这超出了任何单一自动指标所能衡量的范围。我们研究了来自预训练音频语言模型(ALM)的结构化、基于评分标准的奖励,作为自回归和扩散式音乐生成器的训练信号。ALM根据评分标准对每个生成的片段进行评分;我们对同一文本提示生成的不同候选片段按其分数进行排名,并将这些排名转换为偏好对,用于在MusicGen-small和ACE-Step~v1上进行DPO(直接偏好优化),此外,我们还将评分标准分数直接用作ACE-Step~v1上DiffusionNFT的标量奖励。在MusicCaps上,基于评分标准的优化同时提升了CLAP、SongEval和Audiobox-Aesthetics的得分,其中DiffusionNFT在ACE-Step上取得了最大的提升。相比之下,在MusicGen-small上,从这些自动评估器中的任何一个构建偏好都会产生明显的跨指标权衡:目标评估器得到改善,而其他独立评估器则出现下降。我们进一步研究了节奏、调性和配器,在这些方面存在精确的目标奖励。直接优化这些专门的奖励能够可靠地改善目标属性,而ALM评分标准仅对节奏和配器提供部分迁移,对调性则没有可测量的改善。综合来看,这些结果表明了一种实用的分工:ALM评分标准对于难以形式化的广泛感知质量是有效的,而当存在可靠的测量方法时,专门的客观奖励仍然更可取。

英文摘要

Post-training text-to-music generation requires reward signals that capture multiple aspects of musical quality beyond what any single automatic metric can measure. We study structured, rubric-based rewards from pretrained audio-language models (ALMs) as training signals for both autoregressive and diffusion-based music generators. An ALM scores each generated clip against the rubric; we rank candidates generated for the same text prompt by their scores and convert these rankings into preference pairs for DPO on both MusicGen-small and ACE-Step v1, and additionally use the rubric scores directly as scalar rewards for DiffusionNFT on ACE-Step v1. On MusicCaps, rubric-based optimization improves CLAP, SongEval, and Audiobox-Aesthetics simultaneously, with the strongest gains obtained by DiffusionNFT on ACE-Step. By contrast, on MusicGen-small, building preferences from any one of these automatic evaluators produces clear cross-metric trade-offs: the targeted evaluator improves while other independent evaluators deteriorate. We further study tempo, key, and instrumentation, where precise objective rewards are available. Directly optimizing these specialized rewards reliably improves the target attributes, whereas ALM rubrics provide only partial transfer for tempo and instrumentation and no measurable improvement for key. Together, these results suggest a practical division of labor: ALM rubrics are effective for broad perceptual qualities that are difficult to formalize, while specialized objective rewards remain preferable when reliable measurements are available.

发表机构

  • University of Washington(华盛顿大学)
  • Paul G. Allen School of Computer Science & Engineering, University of Washington(华盛顿大学保罗·G·艾伦计算机科学与工程学院)
  • Allen Institute for AI(艾伦人工智能研究所)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑