arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

共享高斯化:高斯正则化器对对比学习的认证及其遗漏之处

Shared Gaussianization: What Gaussian Regularizers Certify About Contrastive Learning, and What They Miss

Ruoyu Zhao, Yuting Chen, Jinheng Zhang, Zhehao Zou, Tong Che

arXiv 2610.10299首次发表:更新:

发表机构

City University of Hong Kong; Georgia Institute of Technology; University of Pennsylvania; The Chinese University of Hong Kong; NVIDIA Research(香港城市大学; 佐治亚理工学院; 宾夕法尼亚大学; 香港中文大学; 英伟达研究院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文研究共享高斯化(SG)正则化器对对比学习的认证能力,证明其能检测错位与非均匀性,并给出InfoNCE超出量的平方根速率界,同时揭示纯SG在干扰通道下的局限及优化差距。

AI 中文摘要

像LeJEPA中的SIGReg这样的分布匹配正则化器能对对比学习认证什么?我们研究了共享高斯化(SG),这是一种对两个归一化视图的平均值进行特征函数高斯性检验的方法,该平均值由独立的$\chi_d$半径缩放。由于不一致的视图会缩短平均值,一个检验就能同时检测错位和非均匀性。SG在总体InfoNCE的对齐、均匀最小化器处恰好消失,并且在边际相等的情况下,它将InfoNCE超出量界定为$4\cdot 3^{3/4}\beta$乘以SG损失的平方根,再加上一项与损失线性相关的项。平方根速率和这个无维度常数是尖锐的,并且视图对上的任何平方均值嵌入距离都无法达到更快的速率。通过一个显式的对齐项,一个旋转不变的均匀性检验给出线性界,当且仅当其谱支配InfoNCE核$e^{\beta u^\top v}$的谱时;SG自身的检验满足此条件,高斯核$e^{-\gamma \\|u-v\\|^2}$在$\gamma \ge \beta/2$时恰好符合,而矩匹配永远不满足。在远离最优解时,目标函数有所不同。沿着各向同性的干扰通道,当共享码非均匀时,纯SG通过添加每个视图的干扰来降低其损失。高于通道增益的对齐权重使无干扰解成为严格的局部最小化器;对于LeJEPA,同样的规则给出了一个随批次大小减小而变化的临界SIGReg权重。在有限批次大小下,一个非对角U统计量消除了朝向错位的插件偏差。在受控潜变量模型中,纯SG保留每个视图的风格,高于测量增益的对齐权重会移除它,而对于LeJEPA在三个批次大小下,测量增益区分了保留风格的编码器和不保留风格的编码器。InfoNCE训练也达到了比从头开始的SG$_{0.2}$训练更低的SG$_{0.2}$损失,这指向了一个优化差距。

英文摘要

What can a distribution-matching regularizer such as SIGReg in LeJEPA certify about contrastive learning? We study shared Gaussianization (SG), a characteristic-function Gaussianity test on the average of two normalized views, scaled by an independent $χ_d$ radius. Because disagreeing views shorten the average, one test detects both misalignment and non-uniformity. SG vanishes exactly at the aligned, uniform minimizers of population InfoNCE, and under equal marginals it bounds the InfoNCE excess by $4\cdot 3^{3/4}β$ times the square root of the SG loss, plus a term linear in the loss. The square-root rate and this dimension-free constant are sharp, and no squared mean-embedding distance on view pairs achieves a faster rate. With an explicit alignment term, a rotation-invariant uniformity test gives a linear bound if and only if its spectrum dominates that of InfoNCE's kernel $e^{βu^\top v}$; SG's own test does, Gaussian kernels $e^{-γ\|u-v\|^2}$ qualify exactly when $γ\ge β/2$, and moment matching never does. Away from the optimum, the objectives differ. Along an isotropic nuisance channel, pure SG lowers its loss by adding per-view nuisance whenever the shared code is non-uniform. An alignment weight above the channel's gain makes the nuisance-free solution a strict local minimizer; for LeJEPA, the same rule gives a critical SIGReg weight that decreases with the batch size. At finite batch size, an off-diagonal U-statistic removes a plug-in bias toward misalignment. In controlled latent-variable models, pure SG retains per-view style, an alignment weight above the measured gain removes it, and for LeJEPA at three batch sizes the measured gain separates the encoders that retain style from those that do not. InfoNCE training also reaches a lower SG$_{0.2}$ loss than SG$_{0.2}$ training from scratch, which points to an optimization gap.

Comments27 pages, 4 figures, 3 tables. Ruoyu Zhao and Yuting Chen contributed equally; Tong Che is the project lead

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑