发表机构
Shiga University(滋贺大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出Composer2Vec,通过作曲家条件Transformer学习符号旋律生成中的作曲家嵌入,发现其形成可解释的连续潜在空间,捕捉音乐历史结构,且优于通用音频-文本嵌入。
AI 中文摘要
我们分析了由作曲家条件Transformer学习的作曲家嵌入,将其视为作曲风格的连续潜在空间,而不仅仅是用于生成的内部表示。一个递归预测旋律延续的模型在从MIDI数据中提取的旋律序列上进行了训练,并以作曲家身份(124位作曲家)为条件。对学习到的作曲家嵌入矩阵(124x128)进行主成分分析表明,第一主成分与作曲家出生年份强相关(r = -0.884,p < 0.001,n = 123),这一相关性比我们将相同的PC1-出生年份分析应用于现有的通用音频-文本嵌入(CLAP、MuQ-MuLan)所获得的相关性更强,这些嵌入是在不相关的音频-文本语料库上训练的,而非符号旋律生成。一个洗牌测试(2,000次排列)确认了风格时期标签的轮廓系数具有统计显著性(0.0110,p < 0.001)。我们进一步表明,嵌入空间中的向量算术捕捉了作曲家之间有意义的风格关系。这些结果表明,仅以作曲家身份为监督信号学习到的作曲家嵌入,形成了一个可解释的潜在空间,捕捉了音乐历史结构。
英文摘要
We analyze the composer embeddings learned by a composer-conditioned Transformer as a continuous latent space of compositional style, rather than merely as an internal representation for generation. A model that recursively predicts melody continuations was trained on melodic sequences extracted from MIDI data, conditioned on composer identity (124 composers). Principal component analysis of the learned composer embedding matrix (124x128) shows that the first principal component correlates strongly with composer birth year (r = -0.884, p < 0.001, n = 123), a stronger correlation than we obtain by applying the same PC1-birth-year analysis to existing general-purpose audio-text embeddings (CLAP, MuQ-MuLan) trained on unrelated audio-text corpora, not on symbolic melody generation. A shuffle test (2,000 permutations) confirms that the Silhouette score for stylistic-period labels is statistically significant (0.0110, p < 0.001). We further show that vector arithmetic in the embedding space captures meaningful stylistic relationships between composers. These results suggest that composer embeddings, learned without supervision beyond composer identity, form an interpretable latent space that captures musical-historical structure.