arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

ViTAMINS:使用合成难负样本训练自监督视觉Transformer的实证研究

ViTAMINS: An Empirical Study of Training Self-Supervised Vision Transformers with Synthetic Hard Negatives

Nikos Giakoumoglou, Andreas Floros, Kleanthis-Marios Papadopoulos, Tania Stathaki

arXiv 2609.01041首次发表:更新:

发表机构

Imperial College London(帝国理工学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究提出ViTAMINS方法,将合成难负样本用于视觉Transformer预训练,经多任务基准测试,该方法性能优于竞争模型且资源效率更高,推动对比学习成为主流生成式与自蒸馏方法的替代方案。

AI 中文摘要

我们提出ViTAMINS,一种将合成难负样本融入无监督视觉Transformer预训练以提升表示质量的方法。该方法在ImageNet及迁移学习、图像检索、复制检测、图像与视频分割任务上经过全面基准测试。值得注意的是,所提出的负样本使学习到的表示呈现出涌现特性,包含图像语义内容的明确信息,且可作为出色的分类器(较基线提升达11.3%)。ViTAMINS通过对现有对比框架进行简单修改实现上述优势,在优于竞争方法的同时资源效率更高,例如其ViT-B模型超越了使用ViT-L的V-JEPA。研究结果促使人们重新将对比学习视为优于主流生成式和自蒸馏方法的更简单且强大的替代方案。

英文摘要

We introduce ViTAMINS, a method that integrates synthetic hard negatives into unsupervised vision transformer pretraining to improve representation quality. Our approach is thoroughly benchmarked on ImageNet and transfer learning, image retrieval, copy detection, and image, video segmentation tasks. Notably, our proposed negatives give rise to emergent properties, where learned representations contain explicit information about the semantic content of an image and serve as excellent classifiers (up to +11.3% over baselines). ViTAMINS achieves these benefits through simple modifications to existing contrastive frameworks and outperforms competing methods while being more resource efficient, e.g., our ViT-B surpasses V-JEPA with ViT-L. Our findings motivate reconsidering contrastive learning as a simpler yet powerful alternative to dominant generative and self-distillation approaches.

CommentsWACV 2027

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑