发表机构
Université Paris-Saclay; CNRS; CentraleSupélec; LISN(巴黎-萨克雷大学; 法国国家科学研究中心; 中央理工-高等电力学院; LISN实验室)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出统一JEPA表示坍缩缓解策略的早期训练稳定性理论,并基于此设计ResidualPred预测器,在表格基准和图像预训练中提升有效秩与下游准确率。
AI 中文摘要
联合嵌入预测架构(JEPAs)容易遭受表示坍缩,通常通过经验性启发式方法来缓解。我们发展了一种早期训练稳定性理论,统一了这些启发式方法。在平凡不动点附近对耦合的JEPA梯度流进行线性化,揭示了两种竞争效应:驱动力($\gamma$)和衰减效应($\sigma$)。在近似谱解耦下,每模态稳定性比率 $\mu_i = \gamma_i / \sigma_i$ 可分解为独立的数据侧和预测器侧项,且不稳定模态的数量追踪了可能出现的表示的秩。该框架预测了一个相边界,我们在超过800个Tabular-JEPA配置中经验性地确认了这一点。它还将预测器缩放、掩蔽比率和EMA统一为改变$\mu$的不同机制。在此分析的指导下,我们引入了ResidualPred,一种在初始化时注意力偏向恒等变换的Transformer预测器;它在表格基准以及CIFAR-10、CIFAR-100、STL-10和ImageNet上的I-JEPA预训练中,提高了有效秩和下游准确率。我们的框架将经验性的坍缩避免启发式方法与明确的动力学图景联系起来,产生了理论驱动的稳定器。代码可在该https URL获取。
英文摘要
Joint-Embedding Predictive Architectures (JEPAs) are prone to representation collapse, typically mitigated through empirical heuristics. We develop an early-training stability theory that unifies these heuristics. Linearising the coupled JEPA gradient flow around the trivial fixed point reveals two competing effects: a driving force ($γ$) and a decay effect ($σ$). Under approximate spectral decoupling, a per-mode stability ratio $μ_i = γ_i / σ_i$ factorises into independent data-side and predictor-side terms and the count of unstable modes tracks the rank of representations that can emerge. The framework predicts a phase boundary, which we confirm empirically across more than 800 Tabular-JEPA configurations. It also unifies predictor scaling, masking ratio, and EMA as distinct mechanisms for shifting $μ$. Guided by this analysis, we introduce ResidualPred, a transformer predictor whose attention is biased toward the identity at initialisation; it improves both effective rank and downstream accuracy on tabular benchmarks and in I-JEPA pretraining on CIFAR-10, CIFAR-100, STL-10, and ImageNet. Our framework connects empirical collapse-avoidance heuristics to an explicit dynamical picture, yielding theory-driven stabilizers. Code is available at https://github.com/jose-melo/drive-vs-decay.
CommentsAccepted at NeurIPS 2026 (Main Track). 57 pages, 10 pages of main text; appendices and the NeurIPS paper checklist included. Code: https://github.com/jose-melo/drive-vs-decay