arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

什么构成了好的层?评估音乐基础模型的逐层固有属性

What Makes a Good Layer? Assessing the Layer-Wise Intrinsic Properties of Music Foundation Models

Angelos-Nikolaos Kanatas, Yuexuan Kong, Pablo Alonso-Jiménez, Xavier Serra, Dmitry Bogdanov

arXiv 2608.14819首次发表:更新:

AI 中文总结

该研究分析12个音乐基础模型的逐层固有属性,发现现有指标无法通用追踪所有音乐任务的层质量,提出音高转置等变度指标,该指标可有效用于层选择,性能优于可训练多层融合方法,尤其在数据有限场景

AI 中文摘要

音乐基础模型通常被用作冻结的音频特征提取器,但选择从哪一层提取特征在很大程度上仍是凭经验进行的。当前的做法默认采用固定深度或多层融合,对于为何某些层在下游任务间的迁移效果更好,以及表示质量如何随深度和预训练范式变化,人们的理解十分有限。我们对12个音乐基础模型开展了系统的逐层分析,这些模型涵盖三种预训练范式(掩码建模、自回归建模和对比学习),通过固有几何属性和基于变换的属性来表征它们的隐藏表示。我们将无标签的表示质量指标与15个下游任务的逐层性能相关联,发现多个指标可追踪流派分类、情感识别、自动标注和节拍跟踪的层质量,尽管在不同任务和预训练范式下强度有所差异。然而,所有指标在调性任务(如调式估计和弦识别)上均失效,表明不存在单一属性可作为音乐信息检索任务间表示质量的通用代理。为解决这一缺口,我们引入了一种音高转置等变度指标,该指标可捕获这些标准指标遗漏的属性,为不同模型家族提供一致的调性质量指标。最后,我们表明固有指标可作为层选择的有效代理,其性能可与可训练的多层融合方法相媲美甚至更优,尤其在数据有限的场景中表现突出。

英文摘要

Music foundation models are commonly used as frozen audio feature extractors, yet selecting which layer to extract from remains largely heuristic. Current practice defaults to fixed depths or multi-layer fusion, with limited understanding of why certain layers transfer better across downstream tasks or how representation quality varies with depth and pre-training paradigm. We conduct a systematic layer-wise analysis of 12 music foundation models spanning three pre-training paradigms (masked modeling, autoregressive modeling, and contrastive learning), characterizing their hidden representations through intrinsic geometric and transformation-based properties. Correlating label-free representation-quality metrics with layer-wise performance across 15 downstream tasks, we find that several metrics track layer quality for genre classification, emotion recognition, automatic tagging, and beat tracking, albeit with varying strength across tasks and pre-training paradigms. However, all metrics fail on tonal tasks such as key estimation and chord recognition, indicating that no single property serves as a general proxy for representation quality across music information retrieval tasks. To address this gap, we introduce a pitch-transposition equivariance measure that captures properties missed by these standard metrics, providing a consistent indicator of tonal quality across model families. Finally, we show that intrinsic metrics can serve as effective proxies for layer selection, matching or outperforming trainable multi-layer fusion methods, particularly in limited-data settings.

Comments11 pages, 2 figures, 2 tables. Accepted at ISMIR 2026. Project page: https://angeloskanatas.github.io/music-fms-layer-eval/

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑