arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

可扩展的补丁级自监督学习

Scalable Patch-Level Self-Supervised Learning

Maximilian Seitzer, Gabriele Trivigno, Antonín Vobecký, Seungeun Yi, Maxime Oquab, Huy V. Vo, Oriane Siméoni, Piotr Bojanowski

arXiv 2610.10013首次发表:更新:

发表机构

Meta FAIR(Meta FAIR)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

提出JEM,一种原理性的学生-教师自监督学习方法,通过跨视图对齐补丁表征并正则化信息与结构保持,在3亿至70亿参数规模稳定训练,分割性能超越DINOv2和DINOv3。

AI 中文摘要

大规模自监督学习(SSL)能够产生强大的视觉表征。然而,大多数可扩展的SSL方法依赖于多种目标和稳定机制的临时组合。我们退一步思考,能否设计一种高性能且原理性的SSL算法。从多视角假设出发,即任务相关内容由不同视角共享的信息所捕获,我们构建了一个信息论目标函数,并将其分解为可解释的项。该推导产生了JEM,一种学生-教师方法,通过跨视角对齐对应的补丁表征来学习,并显式地通过信息和结构保持损失进行正则化。JEM在3亿到70亿参数范围内稳定训练,据我们所知,是首个在70亿规模下展示的潜在空间补丁级方法。在所有规模下,JEM在全局和密集探测任务上均达到强性能,在分割基准上持续超越DINOv2算法,而DINOv2是当今最强视觉SSL方法的重要基础。值得注意的是,在70亿参数下,尽管训练数据少12倍且无细化阶段,JEM在全景分割上超越了DINOv3的性能。这些结果表明,我们确实可以设计一种能够学习强表征、原理性且稳定的SSL算法。

英文摘要

Self-supervised learning (SSL) at scale produces powerful visual representations. However, most scalable SSL methods rely on ad hoc combinations of multiple objectives and stabilization mechanisms. Taking a step back, we ask if we can design a high-performing, yet principled SSL algorithm. Starting from the multi-view assumption, stipulating that task-relevant content is captured by the information common to different views, we construct an information-theoretic objective decomposing into interpretable terms. This derivation yields JEM, a student-teacher method that learns by aligning corresponding patch representations across views, explicitly regularized by information and structure preservation losses. JEM trains stably from 300M to 7B parameters, and, to our knowledge, is the first latent-space patch-level method demonstrated at 7B scale. Across all scales, JEM reaches strong performance on both global and dense probing tasks, on segmentation benchmarks consistently surpassing the DINOv2 algorithm, an influential foundation for today's strongest visual SSL methods. Notably, at 7B parameters, it exceeds the performance of DINOv3 on panoptic segmentation, despite being trained on $12\times$ less data without refinement stages. These results demonstrate that we can indeed design an SSL algorithm that learns strong representations, is principled and stable.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑