AI 中文总结
本研究提出CoJEPA,将对比学习与JEPA结合,在MIR任务上性能优于或匹配单独方法,无需额外参数或任务特定架构改动,鼓励关注智能训练目标而非增大模型规模。
AI 中文摘要
联合嵌入预测架构(Joint-Embedding Predictive Architecture,JEPA)通过隐空间自监督预测学习丰富表征的性能已得到验证,但它通常依赖带指数移动平均(EMA)的教师-学生架构来稳定训练,且易产生无信息的表征。对比学习训练稳定,能生成强全局表征,但因目标的全局特性,在局部任务上表现受限。本研究将二者结合为CoJEPA:一个单一共享主干,同时在掩码序列令牌上采用JEPA目标训练,在类别令牌上采用对比目标训练。对比梯度提供稳定性,完全无需EMA教师;JEPA则通过对比学习无法提供的局部预测丰富序列令牌。关键在于,主干未增加额外参数,仅通过训练信号设计引导模型生成更丰富的表征。CoJEPA融合了二者优势,在全局与局部音乐信息检索(MIR)任务上的性能优于或匹配两种单独方法,在调性与和声理解上尤其表现突出,且无需任何特定任务的架构改动。CoJEPA表明,结合具有互补归纳偏置的目标可替代模型规模,鼓励未来研究投入更智能的训练目标,而非一味增大模型规模。
英文摘要
Joint-Embedding Predictive Architecture (JEPA) has shown strong performance in learning rich representations through self-supervised prediction in latent space. However, it typically relies on teacher--student architecture with an EMA to stabilise training, and can tend to yield uninformative representations. Contrastive learning is stable to train and produces strong global representations, but remains limited on local tasks by the global nature of its objective. In this work, we combine both into CoJEPA: a single shared backbone jointly trained with a JEPA objective on masked sequence tokens and a contrastive objective on the class token. The contrastive gradient provides stability, removing the need for an EMA teacher entirely, while JEPA enriches the sequence tokens via local predictions that contrastive learning alone cannot provide. Crucially, no extra parameters are added to the backbone: the same model is guided towards richer representations purely through the design of its training signal. CoJEPA takes the best of both worlds, outperforming or matching both individual methods across global and local MIR tasks, with a particularly strong advantage on tonal and harmonic understanding, and without any task-specific architectural changes. CoJEPA shows that combining objectives with complementary inductive biases can substitute for scale, encouraging future work to invest in smarter training objectives over ever-larger models.