AI 中文总结
本研究提出基于LeJEPA自监督预训练的仅注意力白盒Transformer,减少约31%参数,在CIFAR-10、CIFAR-100上实现与原模型相当的准确率,还发现标准ViT中MLP模块存在冗余。
AI 中文摘要
现有针对白盒网络的自监督学习研究,通常通过优化算法推导白盒网络,与自监督学习范式相分离。本研究从联合视角重新审视这两个组件。基于LeJEPA的自监督框架假设,下游任务的最优嵌入分布为各向同性高斯分布,这在概念上等价于指导白盒Transformer优化的稀疏率减少目标中的扩展项R(Z)。基于该观察,本研究使用LeJEPA自监督范式优化R(Z),并通过交替方向乘子法(ADMM)将剩余项R^c(Z|U_[K])+λ||Z||₀推导为仅注意力Transformer,摒弃了原始设计的ISTA结构或MLP层。实验结果表明,在LeJEPA自监督范式下,本研究的仅注意力白盒Transformer在Base规模下,于CIFAR-10数据集上达到88.88%的分类准确率,于CIFAR-100数据集上达到63.54%的分类准确率;而原始白盒Transformer CRATE在CIFAR-10上达到89.18%的准确率,在CIFAR-100上达到63.56%的准确率。本模型在实现有竞争力性能的同时,减少了约31%的参数数量。除白盒设置外,本研究还探究了标准ViTs,发现在知识蒸馏下将所有MLP块替换为ReLU激活函数,可减少约66%的参数,同时保留有竞争力的准确率,为进一步探究标准ViT架构中MLP模块的冗余性提供了动机。
英文摘要
Existing studies on self-supervised learning for white-box networks typically decouple the derivation of white-box networks via optimization algorithms from self-supervised learning paradigms. In this work, we instead revisit the two components from a joint perspective. The LeJEPA-based self-supervised framework assumes an isotropic Gaussian distribution as the optimal embedding distribution for downstream tasks, which is conceptually equivalent to the expansion term $R(Z)$ in the sparse rate reduction objective guiding white-box Transformer optimization. Building on this observation, we use the LeJEPA self-supervised paradigm to optimize $R(Z)$, and derive the remaining terms $R^{c}(Z\mid U_{[K]})+λ\lVert Z\rVert_{0}$ via the alternating direction method of multipliers (ADMM) into an attention-only Transformer that dispenses with the ISTA structure or MLP layers of the original design. Experimental results demonstrate that our attention-only white-box Transformer achieves classification accuracies of $88.88\%$ on CIFAR-10 and $63.54\%$ on CIFAR-100 at the Base scale under the LeJEPA self-supervised paradigm, while the original white-box Transformer CRATE achieves classification accuracies of $89.18\%$ on CIFAR-10 and $63.56\%$ on CIFAR-100. Our model achieves competitive performance while reducing the parameter count by roughly $31\%$. Beyond the white-box setting, we further investigate standard ViTs and find that replacing all MLP blocks with ReLU activations under knowledge distillation removes approximately 66\% of the parameters while preserving competitive accuracy, motivating further investigation into the potential redundancy of MLP modules in standard ViT architectures.