扭结与光滑性:拉普拉斯类源分布下实解析nICA的可辨识性
Kinks vs. Smoothness: Identifiability of Real Analytic nICA for Laplace-like Sources
浏览论文内容
中文总结 AI 辅助
本研究证明实解析生成函数在源分布具有有限一阶导数不连续点(如拉普拉斯分布)时可辨识,利用扭结与光滑性对比,适用于归一化流和变分自编码器,并在CelebA数据上恢复可解释潜在因子。
中文摘要 AI 辅助
许多机器学习系统试图用生成复杂数据(如图像或金融时间序列)的隐藏独立因子来解释这些数据。恢复真正的底层因子,而非其某种打乱版本,是非线性独立成分分析(nICA)的核心挑战。我们证明了当源概率密度函数在一阶导数中具有有限数量的不连续点(即“扭结”)时,实解析生成函数在平凡歧义范围内是可辨识的(即精确恢复)。拉普拉斯分布是满足该假设的最突出例子。我们的证明依赖于源分布中的扭结与实解析函数光滑性之间的对比。实解析函数构成了一类广泛的生成机制,并且可以用标准激活函数(如tanh、softplus、GELU)的归一化流或变分自编码器进行近似,因此我们的结果几乎无需修改即可应用于现有训练流程。我们在真实和合成数据上使用归一化流和变分自编码器进行了实验,展示了它们的可辨识性。在CelebA数据的实验中,我们恢复了多个可解释的潜在因子,这些因子控制了数据集中的独特属性。
英文摘要
Many machine learning systems try to explain complex data - like images or financial time series - in terms of hidden, independent factors that generated them. Recovering the true underlying factors, rather than some scrambled version of them, is the central challenge of nonlinear Independent Component Analysis (nICA). We prove identifiability (exact recovery) up to trivial ambiguities for real analytic generating functions when source probability density functions have a finite number of discontinuities in the first derivative. The Laplace distribution is the most prominent example satisfying this assumption. Our proof relies on the contrast between kinks in the source distribution and the smoothness of real analytic functions. Real analytic functions comprise a broad class of generating mechanisms, and can be approximated with Normalizing Flows or Variational Autoencoders with standard activation functions (e.g., tanh, softplus, GELU), so our result applies with minimal changes to existing training pipelines. We perform experiments on real and synthetic data with both Normalizing Flows and Variational Auto-Encoders demonstrating their identifiability properties. In experiments on CelebA data we recover several interpretable latent factors controlling unique attributes across the dataset.
发表机构
- University of Florida(佛罗里达大学)
机构由 AI 辅助整理,请以论文原文为准。