AI 中文总结
本文提出ARIMA框架,用于符号音乐的基于重构的预测表示学习。它直接从数据学习紧凑窗口表示,编码窗口为潜在表示,训练因果预测器并通过结构化重构确定编码器基础,在多下游任务表现良好,证明相关方法的有效性和重要性。
AI 中文摘要
符号音乐的自监督学习很大程度上通过令牌级预训练取得了进展,但此类表示仍与特定于分词器的序列相关联,并且通常只能间接提供时间跨度级嵌入。在本文中,我们提出了ARIMA,这是一种用于符号音乐的基于重构的潜在预测框架,它直接从数据中学习紧凑的基于窗口的表示。ARIMA将每个固定持续时间的窗口编码为连续的潜在表示,通过对比下一个潜在预测训练因果预测器,并通过音乐元素的结构化重构来确定编码器的基础。这种设计在对跨窗口的时间进展进行建模时保留了局部音乐细节。我们在跨越各种音乐理解水平的下游任务上评估了ARIMA。结果表明,ARIMA在涉及和声、节奏和跨表演检索的任务上特别高效且有效,而在其他任务上与大得多的基线相比仍具有竞争力。消融实验进一步表明,下一个潜在预测对于时间上集成的表示至关重要,并且结构化重构在不需要显式方差正则化的情况下稳定了潜在学习。代码位于此https URL。
英文摘要
Self-supervised learning for symbolic music has advanced largely through token-level pretraining, but such representations remain tied to tokenizer-specific sequences and often provide time-span-level embeddings only indirectly. In this paper, we propose ARIMA, a reconstruction-grounded latent predictive framework for symbolic music that learns compact window-based representations directly from data. ARIMA encodes each fixed-duration window into a continuous latent representation, trains a causal predictor with contrastive next-latent prediction, and grounds the encoder through structured reconstruction of music elements. This design preserves local musical details while modeling temporal progression across windows. We evaluate ARIMA on downstream tasks spanning various levels of music understanding. Results show that ARIMA is particularly efficient and effective on tasks involving harmonic, timing, and cross-performance retrieval, while remaining competitive with much larger baselines on other tasks. Ablations further show that next-latent prediction is essential for temporally integrated representations, and that structured reconstruction stabilizes latent learning without requiring explicit variance regularization. The code is at https://github.com/AndyWeasley2004/symbolic_music_wm.