流派偏见还是审美感知?识别和减轻音乐评价中的捷径学习
Genre Bias or Aesthetic Perception? Identifying and Mitigating Shortcut Learning in Music Evaluation
浏览论文内容
中文总结 AI 辅助
研究音乐评价模型中流派诱导的捷径学习问题,提出联合重新加权硬样本并规范化组级性能的训练目标,减少流派相关偏差,提高与人类偏好的一致性。
中文摘要 AI 辅助
音乐美学评分在数据集整理、生成模型评估和音乐生成奖励建模等应用中起着关键作用。近期方法依赖于基于人工标注评分训练的深度神经网络,但这些模型可能利用虚假相关性而非捕捉有感知意义的美学。本文识别出音乐评价模型中一种未充分探索的失败模式:流派诱导的捷径学习。通过对SongEval的系统分析表明,训练数据中的偏差导致流派相关特征与预测分数之间存在强相关性,使模型将其用作美学代理。这导致对流行音乐的系统高估和其他流派高质量样本的低估,预测与人类偏好不一致。为解决此问题,我们提出一个联合重新加权硬样本并规范化组级性能的训练目标,鼓励模型学习音乐性的流派不变表示。实验结果表明,我们的方法减少了流派相关偏差,提高了与人类偏好的一致性。
英文摘要
Music aesthetics scoring plays a critical role in applications such as dataset curation, generative model evaluation, and reward modeling for music generation. Recent approaches rely on deep neural networks trained on human-annotated ratings, but these models may exploit spurious correlations rather than capturing perceptually meaningful aesthetics. In this work, we identify a previously underexplored failure mode in music evaluation models: genre-induced shortcut learning. Through a systematic analysis of SongEval, we show that biases in training data lead to strong correlations between genre-related features and predicted scores, causing the model to use them as a proxy for aesthetics. This results in systematic overestimation of pop music and undervaluation of high-quality samples from other genres, leading to predictions that are inconsistent with human preferences. To address this issue, we propose a training objective that jointly reweights hard samples and regularizes group-level performance, encouraging the model to learn genre-invariant representations of musicality. Experimental results demonstrate that our method reduces genre-dependent bias and improves alignment with human preferences, as reflected by gains in both cross-genre and within-genre preference alignment.