arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

两种全局裁剪足矣:定位DINO式自监督学习中的语义涌现

Two Global Crops Suffice: Locating Semantic Emergence in DINO-Style Self-Supervised Learning

Basavaraj Sunagad, Artur Jesslen, Adam Kortylewski

arXiv 2609.28187首次发表:更新:

发表机构

CISPA Helmholtz Center for Information Security; University of Freiburg(CISPA亥姆霍兹信息安全中心; 弗莱堡大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究通过实证剖析DINO式自监督学习,发现语义涌现主要源于同一实例的全局视图对齐,iBOT目标仅起细化作用,并指出语义对应比分类准确率更能预测下游性能。

AI 中文摘要

使用DINO式目标训练的自我监督视觉变换器在各种视觉任务中表现出惊人的涌现语义表示质量,然而其背后的机制仍不清楚。我们对DINO家族进行了系统的实证剖析,并表明语义表示主要源于对同一图像实例的几何不同全局视图之间一致性的强制约束。这种实例特定的全局对齐充当了DINO式学习的语义锚点。在受控重训练实验中,我们在语义对应以及一系列2D和3D下游任务上评估发现,仅当与这种全局对齐联合训练时,补丁级掩蔽目标才能增强语义,这表明iBOT目标是对现有语义结构进行细化和致密化,而非独立创造语义。相比之下,在固定计算量下,局部到全局视图对齐相比纯粹的全局对齐并不能实质性地改善语义质量。超越训练设计,我们重新审视了语义表示质量的评估方式:虽然分类准确率是标准验证分数,但语义对应提供了互补的维度,能更可靠地预测下游任务性能。总之,这些发现提供了DINO式学习的功能分解,并代表着向理解语义表示如何在自监督视觉模型中涌现迈出的重要一步。

英文摘要

Self-supervised vision transformers trained with DINO-style objectives exhibit striking emergent semantic representation quality across visual tasks, yet the mechanisms underlying this behavior remain unclear. We present a systematic empirical dissection of the DINO family and show that semantic representations arise primarily from enforcing consistency between geometrically distinct global views of the same image instance. This instance-specific global alignment acts as the semantic anchor of DINO-style learning. Across controlled retraining experiments evaluated on semantic correspondence and a diverse suite of 2D and 3D downstream tasks, we find that patch-level masking objectives enhance semantics only when trained jointly with this global alignment, indicating that the iBOT objective refines and densifies existing semantic structure rather than creating it independently. In contrast, local-to-global view alignment does not substantially improve semantic qualities at fixed compute beyond a purely global alignment. Beyond training design, we revisit how semantic representation quality should be evaluated: while classification accuracy is the standard validation score, semantic correspondence provides a complementary axis that more reliably predicts downstream task performance. Together, these findings provide a functional decomposition of DINO-style learning and represent an important step toward understanding how semantic representations emerge in self-supervised vision models.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑