AI 中文总结
本文证明音乐生成模型的内在信号(预测损失、熵及SAE概念)与人类评分高度相关,据此训练轻量模型实现自动音乐评估,并在五个基准上验证其有效性。
AI 中文摘要
当前的音乐生成模型能够产生高质量的音乐,但这种能力是否意味着它们“理解”了其输出的音乐品质,并且这种理解是否与人类评价一致?先前尝试使用生成模型的似然度来评估音乐(这是文本中常用的一种方法)被证明是不成功的,导致研究者转而依赖独立的监督式音乐评估模型。在本文中,我们对此问题给出肯定回答:我们证明模型的内在信号——源自其隐藏表示和预测——与人类评分高度相关。具体而言,我们研究MusicGen并考虑三类特征:(1)预测损失,(2)预测熵,(3)使用稀疏自编码器(SAE)从模型中提取的概念。利用这些特征,我们训练一个轻量级预测模型来估计主观评分。我们分别及组合地评估这些特征。我们假设这些信号与聆听过程相平行:损失和熵的时域和频域结构反映了听者的期望与惊讶,而SAE潜在空间中的梯度方向预测感知质量。在涵盖连续评分和成对偏好的五个人类评估基准上的实验证实了这一假设,其中SAE潜在变量承载了大部分预测信号。
英文摘要
Current music generative models can produce high-quality music, but does this ability imply that they ``understand'' the musical qualities of their outputs, and is that understanding aligned with human evaluation? Previous attempts to use the likelihood of a generative model to evaluate music, an approach commonly used in text, have proven unsuccessful, leading researchers to rely on standalone supervised music evaluation models. In this paper, we answer this question affirmatively: we show that a model's intrinsic signals---derived from its hidden representations and predictions---are strongly correlated with human ratings. In particular, we study MusicGen and consider three types of features: (1) prediction loss, (2) prediction entropy, and (3) concepts extracted from the model using a sparse autoencoder (SAE). Using these features, we train a lightweight prediction model to estimate subjective ratings. We evaluate these features both individually and in combination. We hypothesize that these signals parallel the listening process: the temporal and frequency-domain structure of loss and entropy reflects listeners' expectation and surprise, while gradient directions in SAE latent space predict perceived quality. Experiments on five human-evaluation benchmarks spanning continuous ratings and pairwise preferences confirm this hypothesis, with SAE latents carrying most of the predictive signal.
CommentsAccepted by the 27th International Society for Music Information Retrieval Conference (ISMIR 2026)