arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

投影器在对比自监督学习中的作用:最后一层秩动态驱动表示质量

On the Role of the Projector in Contrastive Self-Supervised Learning: Last-Layer Rank Dynamics Drive Representation Quality

Siladittya Manna, Priyangshu Mandal, Umapada Pal, Saumik Bhattacharya

arXiv 2609.26334首次发表:更新:

发表机构

Hong Kong Baptist University; Indian Institute of Science; Indian Institute of Technology Kharagpur; Indian Statistical Institute(香港浸会大学; 印度科学理工学院; 印度理工学院卡哈拉格普尔分校; 印度统计学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文通过分析投影器和编码器的秩动态,发现秩降低主要发生在最后一层,据此提出针对最后一层的权重正则化策略,在ImageNet100和CIFAR数据集上优于全网络正则化,提升了表示质量。

AI 中文摘要

在自监督对比学习中,表示的维度坍缩是一个长期存在的问题。一种防止这种表示坍缩的显著技术是使用称为投影器的多层感知器网络。在多项工作中,发现投影器对自监督对比预训练任务中学习到的表示质量有重大影响。然而,问题仍然存在:投影器扮演什么角色?假设投影器缓解了维度坍缩,那么在缺少显式多层感知器(MLP)头的情况下,什么阻止基础编码器的终端层发挥投影器的功能?在本工作中,我们旨在通过实证研究和分析,检查投影器及编码器的秩动态,来研究投影器内部发生的情况。通过数学分析,我们观察到秩降低效应主要发生在最后一层。受此见解启发,我们提出了一种专门应用于最后一层的权重正则化策略。我们证明,这种有针对性的方法比在整个网络上应用正交权重正则化(WeRank)取得更好的性能,无论有无投影器。我们的方法在ImageNet100数据集上的SimCLR中,Top-1准确率提高了超过1%,并在CIFAR数据集上持续优于基线SimCLR变体,支持了我们对投影器作用的解释。

英文摘要

The dimensional collapse of representations in self-supervised contrastive learning is an ever-present issue. One notable technique to prevent such a collapse of representations is using a multi-layered perceptron network called Projector. In several works, the projector has been found to heavily influence the quality of representations learned in a self-supervised contrastive pre-training task. However, the question still lingers. What role does the projector play? Assuming the projector mitigates dimensional collapse, what prevents the terminal layer of the base encoder from functioning as the projector in the absence of an explicit multi-layer perceptron (MLP) head? In this work, we intend to study what happens inside the projector by examining the rank dynamics of the same and the encoder through empirical study and analysis. Through mathematical analysis, we observe that the effect of rank reduction predominantly occurs in the last layer. Motivated by this insight, we propose a weight regularization strategy applied specifically to the last layer. We demonstrate that this targeted approach yields better performance than applying orthogonal weight regularization across the entire network (WeRank), both with and without a projector. Our method improves Top-1 accuracy by more than 1% on SimCLR on the ImageNet100 dataset and consistently outperforms baseline SimCLR variants on CIFAR datasets, supporting our interpretation of the projector's role.

CommentsUnder review at Transactions on Machine Learning Research (TMLR)

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑