arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

从孤立特征到轨道:通过多SAE对齐发现音乐概念

From Isolated Feature to Orbits: Discovering Music Concepts via Multi-SAE Alignment

Liwei Lin, Gus Xia

arXiv 2610.01864首次发表:更新:

发表机构

NYU Shanghai; MBZUAI(上海纽约大学; 穆罕默德·本·扎耶德人工智能大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究提出通过多SAE对齐和音高移位归纳偏置,将音乐概念视为结构化轨道而非孤立特征,在两种音乐基础模型上成功恢复和弦、调性等结构,仅需少量锚示例即可解释概念族。

AI 中文摘要

我们如何理解音乐基础模型在内部学到了什么?大多数可解释性方法,如探针和稀疏自编码器(SAEs),侧重于在最小结构假设下识别单个特征。我们认为许多概念更适合被理解为结构化关系,而非孤立特征。这在音乐中尤为突出,因为音调结构在音高和时间空间中组织。例如,和弦或调性等概念自然地表达为结构化集合(如一个和弦的12个移位,或一个调内的全音阶系统),而非孤立特征。在本研究中,我们从特征识别转向基于结构的分析,探究音乐基础模型学习到的内部表征是否以组织化的特征结构形式出现。为此,我们引入了一个框架,利用音高移位作为归纳偏置,通过多视角SAE对齐来诱导有序轨道。具体而言,我们生成音高移位后的输入对,并对其SAE表征进行对齐,以发现与音高相关的结构化特征组。实验结果表明,该方法在两种最先进的音乐基础模型上恢复了对应于和弦、调性和旋律模式的轨道结构,同时仅需最少的锚定(例如,几个锚示例)即可解释整个概念族。

英文摘要

How can we understand what a music foundation model has learned \textit{internally}? Most interpretability approaches, such as probing and Sparse Autoencoders (SAEs), focus on identifying individual features with minimal structural assumptions. We argue that many concepts are better understood as \textit{structured relations} rather than isolated features. This is especially prominent in music, where tonal structures are organized in the space of pitch and time. For example, concepts such as chords or keys are naturally expressed as structured sets (e.g., the 12 transpositions of a chord or the diatonic system within a key), rather than isolated features. In this study, \textbf{we shift from feature identification to structure-based analysis}, asking whether the learned inner representations of music foundation model emerge as organized structures over features. To this end, we introduce a framework that uses pitch transposition as an inductive bias to induce ordered orbits via multi-view SAE alignment. Concretely, we generate pitch-shifted input pairs and align their SAE representations to discover structured groups of pitch-related features. Experimental results show that this approach recovers orbit structures corresponding to chords, keys, and melodic patterns across two state-of-the-art music foundation models, while requiring only minimal grounding (e.g., a few anchor examples) to interpret entire concept families.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑