MuseCritic:通过自然语言审美评论学习多维度歌曲奖励
MuseCritic: Learning Multi-Aspect Song Rewards through Natural-Language Aesthetic Critiques
浏览论文内容
中文总结 AI 辅助
本研究提出MUSECRITIC半标量奖励模型,通过两阶段训练生成多维度自然语言评论预测歌曲奖励,在SongEval、Music Arena等数据集上效果优于现有方法,结合GRPO还提升了Muse-0.6B的多项审美指标。
中文摘要 AI 辅助
长篇歌曲生成模型在时长、结构完整性和声学复杂度上持续改进,使得可靠的审美奖励对于将这些模型与人类偏好对齐变得愈发重要。然而,针对完整歌曲的奖励模型仍然有限,现有评估器通常在单次前向传播中预测分数,却不提供可读的解释。我们提出MUSECRITIC,一种半标量奖励模型,它生成涵盖五个审美维度的自然语言评论,并将其作为中间表示来预测连续奖励分数。MUSECRITIC遵循两阶段训练流程:教师模型首先为监督微调提供高质量评论,之后微调后的模型生成自身的评论用于奖励学习,以此缓解训练与推理间的分布偏移。在包含200首SongEval歌曲的域内测试集上,MUSECRITIC将宏平均均方误差从0.2875降至0.2316,并将宏平均LCC、SRCC和Kendall's tau分别提升至0.9068、0.8838和0.7178。在包含733个偏好对的域外Music Arena基准上,它达到了71.35%的最高准确率。此外,结合MUSECRITIC与GRPO,Muse-0.6B在SongEval和Audiobox Aesthetics的全部9项审美指标上均有所提升。这些结果表明,基于评论的奖励建模可降低评分误差,并为歌曲生成提供有效的优化信号。项目代码仓库可在该https链接获取。
英文摘要
Long-form song generation models continue to improve in duration, structural coherence, and acoustic complexity, increasing the need for reliable aesthetic rewards aligned with human preferences. However, reward models for complete songs remain limited, and existing evaluators typically predict scores in a single forward pass without readable explanations. To this end, we introduce MuseCritic, a semi-scalar reward model that generates a natural-language critique covering five aesthetic dimensions and uses it as an intermediate representation to predict continuous reward scores. MuseCritic follows a two-stage training pipeline: a teacher model first provides high-quality critiques for supervised fine-tuning, then the fine-tuned model generates its own critiques for reward learning, mitigating training-inference distribution shift. On an in-domain test set of 200 SongEval songs, MuseCritic reduces macro-averaged mean squared error from 0.2875 to 0.2316 and improves macro-averaged LCC, SRCC, and Kendall's tau to 0.9068, 0.8838, and 0.7178, respectively. On the out-of-domain Music Arena benchmark with 733 preference pairs, it achieves 71.35% accuracy and remains competitive with strong music-specific reward models. Using MuseCritic with GRPO also improves Muse-0.6B on all nine aesthetic metrics from SongEval and Audiobox Aesthetics. These results show that critique-conditioned reward modeling reduces scoring error and provides an effective optimization signal for song generation. The project repository is available at https://github.com/WuqnEl/MuseCritic.
发表机构
- Fudan University(复旦大学)
机构由 AI 辅助整理,请以论文原文为准。