arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

JEPA预测器:一种用于遮挡特征补全的可迁移算子

The JEPA Predictor: A Transferable Operator for Occluded Feature Completion

William Nguyen, Christopher Nguyen

arXiv 2607.16274首次发表:更新:

发表机构

Aitomatic, Inc.(自动公司)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究JEPA预测器,通过特征空间线性投影将其连接到非JEPA主机,证明其可跨编码器家族移植,能补全遮挡特征,在不同数据集和掩码比例下提升准确率,无需重新训练模型。

AI 中文摘要

联合嵌入预测架构(JEPAs)在训练时会将预测器与编码器一起训练,但在下游部署时会丢弃预测器,仅从编码器读取特征。预测器本质上是一个从可见上下文特征到掩码位置特征的学习算子,这正是部分视图分类器所需的结构。研究表明该算子可在不同编码器家族间移植。首先发现,在强掩码情况下,在JEPA编码器上保留冻结的预测器可大幅缩小与最强非JEPA判别基线的准确率差距。然后通过特征空间间的单个线性投影,将I-JEPA和V-JEPA 2的冻结预测器连接到四个非JEPA主机(CLIP、DINOv3、DINOv2、MAE)上,并在500张ImageNet-1k图像上进行闭式拟合。在ImageNet-9和斯坦福犬数据集上,跨越三个掩码比例,每个主机-供体对中,相对于每个主机的掩码编码器基线的提升随掩码比例K单调增加。CLIP与I-JEPA预测器配对,在严重遮挡时可恢复ImageNet-9上因掩码而损失的大部分准确率,并将细粒度的斯坦福犬准确率从15.9%提升至52.1%(提高36个百分点)。该机制是可识别的:投影在可见补丁上付出固定成本,预测器在掩码补丁上提供不断增加的收益;在严重遮挡情况下收益占主导。在细粒度分类的低K值时,投影成本超过收益,这定义了线性桥接失效的边界。冻结的JEPA预测器可作为跨编码器家族的遮挡特征补全的可迁移算子,在为每个掩码比例拟合匹配的线性探针时,无需对任何一个模型进行重新训练。

英文摘要

Joint-Embedding Predictive Architectures (JEPAs) train a predictor jointly with their encoder, but downstream deployment discards the predictor and reads features from the encoder alone. The predictor is, by construction, a learned operator from visible-context features to features at masked positions, the structure a partial-view classifier needs. We show that this operator is portable across encoder families. We first establish that, at heavy mask, retaining the frozen predictor on a JEPA encoder substantially closes the accuracy gap against the strongest non-JEPA discriminative baselines. We then bolt the frozen predictors of I-JEPA and V-JEPA 2 onto four non-JEPA hosts (CLIP, DINOv3, DINOv2, MAE) through a single linear projection between feature spaces, fit in closed form on 500 ImageNet-1k images. Across both ImageNet-9 and Stanford Dogs and across three mask fractions, the lift over each host's masked-encoder baseline grows monotonically with the mask fraction K in every host-donor pair. CLIP paired with the I-JEPA predictor recovers most of the accuracy that masking removed on ImageNet-9 at heavy occlusion, and lifts fine-grained Stanford Dogs from 15.9% to 52.1% (+36 pp). The mechanism is identifiable: the projection pays a fixed cost on visible patches and the predictor provides a growing benefit on masked patches; the benefit dominates the heavy-occlusion regime. At low K on fine-grained classification the projection cost exceeds the benefit, defining the boundary where the linear bridge breaks down. The frozen JEPA predictor functions as a portable operator for occluded feature completion across encoder families, requiring no retraining of either model while fitting matched linear probes per mask fraction.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑