AI 中文总结
研究自监督学习模型易产生虚假偏差问题,提出无偏开放世界正则化框架UOWReg,从全局目标转向条件目标保证表示与属性统计独立,经实验验证能减轻偏差、保持精度,还在新任务中防止子群体崩溃。
AI 中文摘要
尽管近期有进展,但自监督学习(SSL)模型和联合嵌入预测架构(JEPAs)仍易在数据集中学习到虚假偏差。这些技术依赖正则化,通过强制全局目标分布防止表示崩溃。然而全局约束不足以防止偏差纠缠。近期方法是条件分布匹配的部分近似。为此提出无偏开放世界正则化(UOWReg)框架,从全局目标转向条件目标保证表示与目标属性统计独立。通过高斯和球形潜在空间验证,条件匹配减轻偏差,在CelebA基准上减少均等赔率违规,保持分类精度。还引入合成雕刻任务,UOWReg有效防止标准SSL中的子群体崩溃。
英文摘要
Despite recent advances, self-supervised learning (SSL) models and Joint-Embedding Predictive Architectures (JEPAs) remain susceptible to learning spurious biases in the dataset. These techniques rely on regularization, which prevents representation collapse by enforcing a global target distribution such as a multivariate Gaussian or a uniform distribution on the sphere. However, these global constraints are insufficient to prevent bias entanglement, as task-irrelevant features can still segregate the latent space into distinct sub-regions. While recent approaches like Entangling and Disentangling (EnD) and Fair Supervised Contrastive Learning (FSCL) empirically debias the latent space, we show that they act as partial approximations of conditional distribution matching. To enforce this matching explicitly, we propose Unbiased Open World Regularization (UOWReg), an encoder-only framework. We show that this shift from a global to a conditional objective guarantees statistical independence between the learned representations and the targeted attributes, regardless of the chosen target distribution. We empirically validate this framework across both Gaussian and spherical latent spaces, using statistical measures to enforce these target distributions. While conditional matching successfully mitigates bias with both distributions, we demonstrate that enforcing conditional uniformity on the sphere yields a lower linearprobing classification error. Empirically, UOWReg reduces Equalized Odds violations on the CelebA benchmark while maintaining competitive classification accuracy compared to existing encoder-only baselines. Furthermore, we introduce the Synthetic Engraving Task-a novel setting in which a dominant macro-structure masks a fine-grained micro-signature. We show that UOWReg effectively prevents the subpopulation collapse observed in standard SSL, successfully isolating micro-signatures even when heavily entangled with the global structure.