arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

用于公平自监督学习的无偏开放世界正则化

Unbiased Open World Regularization for Fair Self-Supervised Learning

L{é}o Nicollier, Marc Pic, Pablo Mus{é}, Enric Meinhardt-Llopis, Gabriele Facciolo

arXiv 2607.22149首次发表:更新:

AI 中文总结

研究自监督学习模型易产生虚假偏差问题,提出无偏开放世界正则化框架UOWReg,从全局目标转向条件目标保证表示与属性统计独立,经实验验证能减轻偏差、保持精度,还在新任务中防止子群体崩溃。

AI 中文摘要

尽管近期有进展,但自监督学习(SSL)模型和联合嵌入预测架构(JEPAs)仍易在数据集中学习到虚假偏差。这些技术依赖正则化,通过强制全局目标分布防止表示崩溃。然而全局约束不足以防止偏差纠缠。近期方法是条件分布匹配的部分近似。为此提出无偏开放世界正则化(UOWReg)框架,从全局目标转向条件目标保证表示与目标属性统计独立。通过高斯和球形潜在空间验证,条件匹配减轻偏差,在CelebA基准上减少均等赔率违规,保持分类精度。还引入合成雕刻任务,UOWReg有效防止标准SSL中的子群体崩溃。

英文摘要

Despite recent advances, self-supervised learning (SSL) models and Joint-Embedding Predictive Architectures (JEPAs) remain susceptible to learning spurious biases in the dataset. These techniques rely on regularization, which prevents representation collapse by enforcing a global target distribution such as a multivariate Gaussian or a uniform distribution on the sphere. However, these global constraints are insufficient to prevent bias entanglement, as task-irrelevant features can still segregate the latent space into distinct sub-regions. While recent approaches like Entangling and Disentangling (EnD) and Fair Supervised Contrastive Learning (FSCL) empirically debias the latent space, we show that they act as partial approximations of conditional distribution matching. To enforce this matching explicitly, we propose Unbiased Open World Regularization (UOWReg), an encoder-only framework. We show that this shift from a global to a conditional objective guarantees statistical independence between the learned representations and the targeted attributes, regardless of the chosen target distribution. We empirically validate this framework across both Gaussian and spherical latent spaces, using statistical measures to enforce these target distributions. While conditional matching successfully mitigates bias with both distributions, we demonstrate that enforcing conditional uniformity on the sphere yields a lower linearprobing classification error. Empirically, UOWReg reduces Equalized Odds violations on the CelebA benchmark while maintaining competitive classification accuracy compared to existing encoder-only baselines. Furthermore, we introduce the Synthetic Engraving Task-a novel setting in which a dominant macro-structure masks a fine-grained micro-signature. We show that UOWReg effectively prevents the subpopulation collapse observed in standard SSL, successfully isolating micro-signatures even when heavily entangled with the global structure.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑