arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.37789cs.LGcs.AI

预测性自监督学习可证明地在干扰下识别随机信号

Predictive Self-Supervised Learning Provably Identifies Stochastic Signals under Nuisance

Fabian A. Mikulasch, Friedemann Zenke

首次发表
浏览论文内容

中文总结 AI 辅助

本文证明预测性自监督学习通过预测互信息最大化和潜在分布匹配两个原则,能在干扰下识别随机信号,并在模拟中验证了其可识别性。

中文摘要 AI 辅助

通过在潜在空间中进行预测而不生成输入数据本身的自监督学习(SSL),能够学习高度抽象且有用的表示。直观上,这种成功常被归因于其丢弃与预测无关的干扰信息的能力。然而,这引发了一个难题:预测相关的潜在信号中的随机变化与真正的干扰都使得观测部分不可预测;它们如何被区分?令人惊讶的是,我们证明了常见的SSL方法能够精确实现这一点,通过隐式实例化一个具有随机动力学和观测私有干扰的潜在变量模型。我们将其恢复随机信号的能力追溯至两个互补原则:预测互信息最大化确保表示保留预测所需的信息,而潜在分布匹配则约束这些信息的编码方式,从而使保留的信号可识别。我们在高斯预测器的模拟中确认了这一可识别性结果,即使在动态且充满干扰的环境中,也能恢复真实信号至仿射变换。

英文摘要

Self-supervised learning (SSL) by predicting in latent space, without generating the input data itself, learns highly abstract, useful representations. Intuitively, this success is often attributed to its ability to discard nuisance information that is irrelevant to prediction. However, this poses a conundrum: both stochastic variation in a prediction-relevant latent signal and true nuisance make observations partly unpredictable; how could they be distinguished? Surprisingly, we prove that common SSL methods can achieve exactly this, by implicitly instantiating a latent-variable model with stochastic dynamics and observation-private nuisance. We trace their ability to recover the stochastic signal to two complementary principles: Predictive mutual information maximization ensures that representations retain the information needed for prediction, while latent distribution matching constrains how this information is encoded, thereby making the retained signal identifiable. We confirm this identifiability result in simulations for Gaussian predictors, which recover the true signal up to an affine transformation even in dynamic, nuisance-laden environments.

发表机构

  • Friedrich Miescher Institute for Biomedical Research(弗里德里希·米舍尔生物医学研究所)
  • University of Basel(巴塞尔大学)

机构由 AI 辅助整理,请以论文原文为准。

↑