AI 中文总结
研究针对符号音乐表示自监督方法探索不足的问题,提出MIDI-RAE-JEPA,结合音高和时间移位等方差目标、LeJEPA及Swin Transformer V2编码器学习分层表示,经实验验证该方法在多方面表现良好,为符号音乐表示提供了可行途径。
AI 中文摘要
丰富的音乐结构内部表示对于诸如机器辅助音乐共同创作等音乐理解任务至关重要,然而符号音乐表示的自监督方法仍未得到充分探索,尤其是那些编码音乐结构分层多尺度性质的方法。我们提出了MIDI-RAE-JEPA,它将音高和时间移位等方差目标与LeJEPA以及Swin Transformer V2编码器相结合,以学习编码为钢琴卷帘图像的符号音乐的分层表示。时间移位等方差目标促使模型内化时间音乐关系。编码器仅通过自监督目标进行训练,包括掩码嵌入预测器,并通过SIGReg防止坍缩。在冻结的编码器嵌入上训练的单独解码器实现了0.995的重建F1,基于这些嵌入的流匹配生成模型产生的生成结果与条件摘录的音高登记和节奏密度紧密匹配,而不匹配的条件产生不相关但音乐上合理的输出。在下游情感分类任务中,学习到的表示优于Haar散射变换基线,并且嵌入距离随音高和时间移位幅度单调增加,证实了可测量的等方差。这些结果表明,基于等方差的自监督学习目标与足够的精细级编码器容量相结合,为符号音乐的语义丰富、生成有用的表示提供了一条可行的途径。
英文摘要
Rich internal representations of musical structure are essential for music understanding tasks such as machine-assisted music co-writing, yet self-supervised approaches for symbolic music representation remain underexplored, particularly those that encode the hierarchical multiscale nature of musical structures. We present MIDI-RAE-JEPA, combining a pitch- and time-shift equivariance objective with LeJEPA and a Swin Transformer V2 encoder to learn such hierarchical representations of symbolic music encoded as piano roll images. The time-shift equivariance objective encourages the model to internalize temporal musical relationships. The encoder is trained purely on self-supervised objectives -- including a masked embedding predictor (MEP) -- with collapse prevented via SIGReg. A separate decoder trained on the frozen encoder embeddings achieves reconstruction F1 of 0.995, and a flow matching generative model conditioned on those embeddings produces generations that closely match the pitch register and rhythmic density of the conditioning excerpt, while mismatched conditioning yields unrelated but musically plausible output. Learned representations outperform a Haar scattering transform baseline on a downstream emotion classification task, and embedding distances increase monotonically with pitch and time shift magnitude, confirming measurable equivariance. These results suggest that equivariance-based SSL objectives, combined with sufficient fine-level encoder capacity, provide a viable path toward semantically rich, generatively useful representations of symbolic music.
Comments8 pages, 8 figures