arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

CI-JEPA:自监督学习中联合嵌入预测架构潜在表征的反事实分析

CI-JEPA: A Counterfactual Analysis of Latent Representations in Joint-Embedding Predictive Architectures for Self-Supervised Learning

Mintu Dutta, Ritesh Vyas, Mohendra Roy *

arXiv 2610.05043首次发表:更新:

发表机构

Pandit Deendayal Energy University; Birla Institute of Technology & Science, Pilani(潘迪特·丁达亚尔能源大学; 比拉理工学院与科学中心(皮拉尼校区))

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

CI-JEPA通过反事实干预感知扩展I-JEPA,预测表征变化并评估选择性敏感性,在Flowers102上以78.14%准确率小幅超越基线,为研究干预诱导的表征变化提供框架。

AI 中文摘要

自监督视觉表征学习在表征训练期间无需人工标注即可学习有用的特征。基于图像的联合嵌入预测架构(I-JEPA)预测被掩码图像区域的潜在表征,但其目标函数并未显式建模对指定视觉干预的响应。我们提出CI-JEPA,一种反事实干预感知的扩展方法,学习预测原始图像与其修改后对应图像之间的表征变化ΔZ = Z_CF - Z。我们通过选择性敏感性来评估表征鲁棒性:对任务相关语义变化的响应强于对干扰变化的响应。在Flowers102上的实验将花朵中心遮挡作为候选语义干预,将背景模糊和背景色调作为候选干扰干预。使用冻结编码器线性探针,CI-JEPA达到78.14%的最佳验证准确率,而预训练的ViT-B/16和I-JEPA基线均为77.55%,提升了0.59个百分点。报告的平均L2表征变化为:中心遮挡4.42,背景色调3.48,背景模糊2.83。这一排序与所评估干预的相对语义选择性一致,而非完全的干扰不变性。准确率比较是补充性的,并未确立相对于基线的鲁棒性提升。这些受控图像修改为研究JEPA表征中干预诱导的变化提供了框架;它们并未确立因果特征发现或对所有视觉变化的鲁棒性。

英文摘要

Self-supervised visual representation learning learns useful features without manual annotations during representation training. The image-based joint-embedding predictive architecture (I-JEPA) predicts latent representations of masked image regions, but its objective does not explicitly model responses to specified visual interventions. We introduce CI-JEPA, a counterfactual intervention-aware extension that learns to predict the representation change $ΔZ = Z_{\mathrm{CF}} - Z$ between an original image and a modified counterpart. We assess representation robustness through selective sensitivity: stronger responses to task-relevant semantic changes than to nuisance changes. Experiments on Flowers102 use flower-center occlusion as a candidate semantic intervention and background blur and tint as candidate nuisance interventions. With frozen-encoder linear probing, CI-JEPA achieves a best validation accuracy of 78.14\%, compared with 77.55\% for both the pretrained ViT-B/16 and the I-JEPA baseline, a gain of 0.59 percentage points. The reported mean $L_2$ representation changes are 4.42 for center occlusion, 3.48 for background tint, and 2.83 for background blur. This ordering is consistent with relative semantic selectivity for the evaluated interventions, rather than complete nuisance invariance. The accuracy comparison is complementary and does not establish improved robustness over the baselines. These controlled image modifications provide a framework for studying intervention-induced changes in JEPA representations; they do not establish causal feature discovery or robustness to all visual changes.

CommentsThis article has been accepted for the International Symposium on Smart Systems, Algorithms & Applications 2026 (IS3A3 2026), Tezpur Central University, India

Journal refInternational Symposium on Smart Systems, Algorithms & Applications 2026 (IS3A3 2026), Tezpur Central University, India

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑