arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.04044cs.CVstat.ML

暹罗JEPA:关于暹罗学生编码器在联合嵌入预测架构中的作用

SiamJEPA: On the Role of Siamese Student Encoders in JEPA

Makoto Yamada

首次发表
浏览论文内容

中文总结 AI 辅助

研究暹罗学生编码器在基于联合嵌入预测架构(JEPA)的表征学习中的作用,提出SiamJEPA,通过实验证明其能有效正则化JEPA目标,提升表征可分性并加速训练,优于单编码器变体。

中文摘要 AI 辅助

最近,联合嵌入预测架构(JEPAs)作为一种有前途的自监督表征学习框架,在计算机视觉和机器学习社区中引起了广泛关注。与重建像素的掩码自动编码器不同,JEPA模型通过预测掩码区域的潜在嵌入来学习表征。现有基于JEPA的方法,如I-JEPA和V-JEPA,通常在学生网络中使用单个编码器。相比之下,在学生网络中使用暹罗编码器更自然地符合受大脑启发的表征学习框架,但其在JEPA模型中的作用在很大程度上仍未得到探索。在本文中,我们研究了暹罗学生编码器在基于JEPA的表征学习中的作用。为此,我们提出了SiamJEPA,即配备指数移动平均(EMA)教师网络的掩码暹罗学生编码器。SiamJEPA也可以看作是受大脑启发的表征学习模型PhiNet的JEPA形式。通过在ImageNet线性探测上的大量实验,我们证明了暹罗编码器作为JEPA目标的有效正则化器,在训练的早期阶段提高了表征的可分性并加速了学习。此外,在有限的训练预算下,SiamJEPA始终优于可比的单编码器JEPA变体,并且比需要更长训练时间的掩码自动编码器(MAE)实现更高的线性探测精度。我们的发现表明,暹罗学生编码器不仅是一种架构选择,而且构成了预测表征学习的重要归纳偏差。这些结果为基于JEPA的模型设计提供了新的见解,并表明纳入暹罗学生架构提供了一种简单而有效的方法来改进自监督表征学习。

英文摘要

Joint Embedding Predictive Architectures (JEPAs) have emerged as a promising framework for self-supervised representation learning by predicting latent embeddings of masked regions rather than reconstructing pixels. Existing JEPA methods typically employ a single student encoder, leaving the role of Siamese student encoders largely unexplored. In this paper, we propose Siamese JEPA (SiamJEPA), a JEPA framework with masked Siamese student encoders and an exponential moving average (EMA) teacher, which can also be viewed as a JEPA formulation of the brain-inspired representation learning model PhiNet. We further introduce Random Shuffle Teacher (RST), which removes spatial correspondence in teacher targets to encourage semantic patch representations, and develop an RST-based semantic-to-spatial curriculum. Experiments on ImageNet show that Siamese student encoders effectively regularize the JEPA objective, improving representation separability and accelerating early-stage learning. Moreover, under RST, stronger Siamese regularization substantially increases class-discriminative information in individual patch tokens, suggesting that the semantic bias arises from the interaction between RST and the Siamese objective rather than from shuffling alone. SiamJEPA consistently outperforms comparable single-encoder JEPA variants under limited training budgets. With RST-based curriculum learning and a ViT-Base backbone, SiamJEPA achieves 74.2\% linear-probing accuracy after 450 epochs using a substantially simpler masking strategy, compared with I-JEPA (72.9\%) and DSeq-JEPA (73.5\%) after 600 epochs. These results demonstrate that Siamese student encoders provide an effective inductive bias for predictive representation learning and can be further enhanced through semantic-to-spatial curriculum learning. The source code is publicly available at https://github.com/oist/SiamJEPA.

发表机构

  • Okinawa Institute of Science and Technology(冲绳科学技术大学院大学)

机构由 AI 辅助整理,请以论文原文为准。

↑