等变音乐Transformer
Equivariant Music Transformer
浏览论文内容
中文总结 AI 辅助
针对标准音乐Transformer等变性不足的问题,提出EMT模型,通过自蒸馏联合优化预测与等变正则化损失,在等变性和生成能力上优于现有方法。
中文摘要 AI 辅助
即使一段音乐在时间上移位或音高上移调,人类仍能识别它,这表明表征空间中存在等变性概念。然而我们的分析显示,标准音乐Transformer会将此类时间移位或音高移调的输入映射到不相关的表征上:随着模型规模扩大或训练时间延长,这些模型的等变性会逐渐降低。这表明标准音乐Transformer中,额外的模型容量被分配用于记忆绝对模式,而非捕捉共享的音乐结构。本文提出等变音乐Transformer(Equivariant Music Transformer,EMT),它通过联合优化下一个token预测任务和辅助等变正则化损失,利用自蒸馏来强制实现等变性。我们发现,额外的等变损失可作为有益的正则化项,同时提升下一个token预测性能并生成等变潜表征。通过客观和主观评估,EMT相较于数据增强、特征工程及最先进(SOTA)基准,展现出更优的等变性和生成能力。更广泛地说,我们的发现表明仅标准语言建模方法无法捕捉音乐的平移对称性,需专用归纳偏置才能生成更好的音乐表征。代码、权重和演示可在线获取。
英文摘要
Humans recognize a musical passage even when it is shifted in time or transposed in pitch, indicating a notion of equivariance in the representation space. Our analysis, however, shows that standard music transformers map such time-shifted or pitch-transposed inputs onto uncorrelated representations: these models become progressively less equivariant as they scale in size or train longer. This suggests that in standard music transformers, additional model capacity is allocated to memorizing absolute patterns rather than capturing shared musical structures. In this paper, we propose the Equivariant Music Transformer (EMT), which enforces equivariance through self-distillation by jointly optimizing a next-token-prediction and an auxiliary equivariance regularization loss. We find that the additional equivariance loss acts as a beneficial regularizer, simultaneously improving next-token prediction and producing equivariant latent representations. Through both objective and subjective evaluations, EMT demonstrates superior equivariance and generative capability compared to data augmentation, feature engineering, and state-of-the-art (SOTA) baselines. More broadly, our findings reveal that standard language modeling methods alone do not capture music's translational symmetries, and dedicated inductive biases are required to produce better music representations. The code, weights and demos are available online.