JEPA 学习掩码无法恢复的内容
JEPA Learns What the Mask Leaves Unrecoverable
中文总结 AI 辅助
本文提出基于小波零空间的测量框架,解释JEPA对掩码类型的敏感性,并通过151次预训练实验验证,指出不可恢复的粗尺度内容是消除捷径的关键。
中文摘要 AI 辅助
联合嵌入预测架构对输入掩码的方式异常敏感:块状掩码有效,分散掩码无效,而现有解释多为经验性的。我们给出一种基于测量的解释。掩码是一种线性测量,在紧支撑小波基下,支撑集完全位于隐藏区域内的每个原子都落在测量的零空间中,在数据中不留痕迹。JEPA 损失仅要求编码上下文足以预测目标,因此低级先验可恢复的目标存在捷径,而移动平均目标编码器可使该捷径自洽。消除捷径的是掩码留下的粗尺度内容无法恢复的部分,前提是每个目标附近有足够的上下文。我们在训练前对该内容进行评分,并在151次预训练运行中检验该解释的独特预测。在ImageNet-100上,条状掩码在面积和连续性上与块状掩码相当,但可恢复,其线性top-1准确率为40.3%,随机掩码为40.8%,而块状掩码为64.3%;在同一几何族内,留下最少不可恢复内容的放置方式比留下最多的低6.5个百分点(五对种子);像素目标跨度7个百分点,而潜在目标跨度25个百分点;与冻结目标相比,随机掩码与块状掩码之间的差距(同类GPU上为19个百分点)缩小至1.5,因此几何效应通过编码器自身产生的目标起作用。在UCF101上,掩码比率决定内容约束还是可达约束起主导作用;移除整帧(在空间上不可恢复,但可从相邻帧恢复)在两种比率下表现最差;在V-JEPA自身的掩码上,完整批处理而非截断变化不大(36.0%对35.1%),而将块内100个目标标记设为可见则提升至48.7%,隐藏块外100个上下文标记则无此效果(33.7%)。
英文摘要
Joint-embedding predictive architectures are unusually sensitive to how the input is masked: block masks work, scattered masks do not, and the explanations are empirical. We give a measurement account. A mask is a linear measurement, and in a compactly supported wavelet basis every atom whose support lies inside the hidden region falls in the measurement's null space and leaves no trace in the data. The JEPA loss asks only that the encoded context suffice for the target, so a target that a low-level prior can recover admits a shortcut, one the moving-average target encoder can make self-consistent. What removes the shortcut is the coarse-scale content the mask leaves unrecoverable, provided enough context stays within reach of each target. We score that content before training and test the account's distinctive predictions in 151 pre-training runs. On ImageNet-100, strip masks match blocks in area and contiguity yet are recoverable, and they land at 40.3% linear top-1, beside random masks at 40.8%, against 64.3% for blocks; within one geometry family, the placements that leave the least unrecoverable content lose 6.5 points to those that leave the most, over five seed pairs; pixel targets span 7 points where latent targets span 25; and against a frozen target the gap between random and block masks, 19 points on the same kind of GPU, closes to 1.5, so the geometry acts through the target the encoder produces for itself. On UCF101 the masking ratio decides which condition, content or reach, binds; removing whole frames, unrecoverable in space but recoverable from neighbouring frames, is worst at both ratios; and on V-JEPA's own masks, batching them intact instead of truncated changes little (36.0% against 35.1%), whereas making 100 target tokens inside the blocks visible lifts them to 48.7% and hiding 100 context tokens outside the blocks does not (33.7%).