发表机构
Vector Institute; University of British Columbia; Carleton University; University of Guelph(向量研究所; 不列颠哥伦比亚大学; 卡尔顿大学; 圭尔夫大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出掩码摇摆者,通过双视图增强与跨视图CLS标记交换改进MAE,在多个任务上显著超越基线,并降低错误率,推动自监督学习发展。
AI 中文摘要
自监督学习(SSL)消除了对标注的需求,并使得模型能够在比监督学习更广泛的领域内发挥作用。自编码器SSL框架通过在瓶颈或噪声注入造成信息丢失后重建自身输入来进行学习。掩码自编码器(MAE)是该框架最成功的实例:它们编码一个随机补丁子集,然后解码被掩码的补丁。在这项工作中,我们引入了关键修改以改进MAE。我们的方法以两种不同方式增强图像,然后分别掩码并编码每个视图。接着,在解码掩码补丁之前,它在视图之间交换全局表示(CLS标记)。通过设计,我们的掩码摇摆者鼓励学习与视图无关的图像摘要,以促进高效迁移。我们进行了大量实验,发现掩码摇摆者在ImageNet-1K kNN上比MAE高出+3-5%,并在细粒度任务上提供了大幅提升,例如在实例检索上相对提升+45%,在动物重识别上+22%,在Omniglot字符识别上+76%。此外,摇摆者在三个新的状态探测数据集上将错误率相对MAE降低了-64%,为世界建模打开了大门。欢迎来到我们的摇摆者派对。
英文摘要
Self-supervised learning (SSL) removes the need for annotations and makes models that are capable across more domains than supervised learning. The autoencoder SSL framework learns by reconstructing its own input after information loss through a bottleneck or noise injection. Masked autoencoders (MAE) are the most successful instantiation of this framework: they encode a random subset of patches, then decode the masked-out patches. In this work, we introduce key modifications to improve MAEs. Our method augments an image in two different ways, then masks and encodes each view separately. It then exchanges the global representations (CLS tokens) between views before decoding the masked patches. By design, our Masked Swingers encourages learning a view-agnostic summary of the image to facilitate efficient transfer. We perform extensive experiments, and find Masked Swingers outperforms MAE by +3-5% on ImageNet-1K kNN and provides large gains on fine-grained tasks, e.g., relative gains of +45% on instance retrieval, +22% on animal re-ID, and +76% on Omniglot character recognition. To boot, Swingers reduces error -64% relative to MAE on three new state-probing datasets, opening the door to world modeling. Welcome to our Swingers party.