发表机构
Imperial College London; CERTH, ITI(帝国理工学院; CERTH ITI)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究提出自监督视觉表征学习的三条必要原则,通过理论证明与受控实验分析各原则的作用,表明主流自监督方法可纳入统一框架,且任意两个原则无法替代第三个。
AI 中文摘要
我们认为,无标签情况下学习视觉表征需要训练信号同时覆盖三个互不重叠的目标:增强视图间的语义不变性、补丁级空间预测以及表征非退化性。我们将这些形式化为观测原则、预测原则和正则化原则,并证明:(i)在无负样本的对齐设置下,结合观测与预测目标而不加入正则化时,恒等编码器是全局极小值点;(ii)这两个目标在编码器输出处是梯度互补的,且结构上无冲突;(iii)动量编码器与在线编码器收敛到同一不动点,且收敛时不提供坍塌保证。对比对齐仅提供自限性的抗坍塌能力,这一点通过显式梯度衰减论证得到形式化。若舍弃预测目标,会因结构设计缺失空间训练信号;若舍弃观测目标,会因结构设计缺失跨视图语义不变性。在我们研究的规模下,任意两个目标都无法替代第三个。所有主流自监督方法均为单一统一能量分解的特例。我们为每一项理论主张都配套了受控实验,包括针对预测目标空间效应的补丁检索评估。
英文摘要
We argue that learning visual representations without labels requires a training signal jointly complete across three non-overlapping objectives: semantic invariance across augmented views, patch-level spatial prediction, and representational non-degeneracy. We formalize these as the observation, prediction, and regularization principles and prove (i) that combining observation and prediction without regularization admits the constant encoder as a global minimizer under negative-free alignment; (ii) that the two objectives are gradient-complementary and structurally non-conflicting at the encoder output; and (iii) that the momentum encoder converges to the same fixed point as the online encoder and provides no collapse guarantee at convergence. Contrastive alignment provides only self-limiting collapse resistance, formalized via an explicit gradient-decay argument. Dropping prediction withholds the spatial training signal by construction; dropping observation forfeits cross-view semantic invariance by construction; at the scale we study, no pair substitutes for the third. Every major self-supervised method is a special case of a single unified energy decomposition. We pair every theoretical claim with a controlled experiment, including a patch-retrieval evaluation for the spatial consequence of prediction.
CommentsECCV 2026 Workshop UniWorld