arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.27853stat.MLcs.LGcs.NEmath.STstat.TH

基于证据学习的变分自编码器的理论研究

Theoretical Study on the Evidential Learning-based Variational Autoencoder

Ge Wang

首次发表
浏览论文内容

中文总结 AI 辅助

本研究证明NIG潜变量层级中重建可见参数空间为三维商空间,通过精确部分最小化前向KL散度实现三坐标约简,并揭示逆典型分配与商状态和先验的显式单调关系。

中文摘要 AI 辅助

正态-逆伽马(NIG)潜变量层级结构具有四个参数,但其诱导的潜变量分布并不能识别全部四个参数。对于$\sigma^2\sim\mathrm{InvGamma}(\alpha,\beta)$,$\mu\mid\sigma^2\sim\mathcal{N}(\gamma,\sigma^2/\nu)$,以及$z\mid\mu,\sigma^2\sim\mathcal{N}(\mu,\sigma^2)$,$z$的边缘分布仅通过$c=\beta(1+1/\nu)$依赖于$(\nu,\beta)$。因此,重建可见的参数空间是三维商空间$(\gamma,\alpha,c)$,并具有一维纤维自由度。对于固定的分层变分目标,对完整NIG先验的前向KL散度进行精确的部分最小化,会在每条纤维上选择一个唯一的先验相对代表,从而得到精确的三坐标约简,其最优值与四坐标目标相同。记$\rho_0=2\beta_0/\nu_0$和$T=c/\{\alpha[(\gamma-\gamma_0)^2+\rho_0]\}$,我们证明逆典型分配$1/\nu_{\rm can}$是$T$的显式严格递增函数。对于$\alpha>1$,比值$u_{\rm epi}/u_{\rm var}=1/\nu_{\rm can}$因此由商状态和先验决定;在共同校准下仅用于排序时,$T$包含相同的逐坐标序数信息。残余先验规范被刻画而非消除:$(\gamma_0,\rho_0)$控制序数依赖性,而$(\nu_0,\alpha_0)$决定数值校准和$1/\nu_{\rm can}$的解析上限。

英文摘要

A normal--inverse-gamma (NIG) latent hierarchy has four parameters, but its induced latent law does not identify all four. For $σ^2\sim\mathrm{InvGamma}(α,β)$, $μ\midσ^2\sim\mathcal{N}(γ,σ^2/ν)$, and $z\midμ,σ^2\sim\mathcal{N}(μ,σ^2)$, the marginal law of $z$ depends on $(ν,β)$ only through $c=β(1+1/ν)$. Hence the reconstruction-visible parameter space is the three-dimensional quotient $(γ,α,c)$, with a one-dimensional fiber degree of freedom. For a fixed hierarchical variational objective, exact partial minimization of the forward KL divergence to a complete NIG prior selects a unique prior-relative representative on each fiber, yielding an exact three-coordinate reduction with the same optimum as the four-coordinate objective. Writing $ρ_0=2β_0/ν_0$ and $T=c/\{α[(γ-γ_0)^2+ρ_0]\}$, we show that inverse canonical allocation $1/ν_{\rm can}$ is an explicit strictly increasing function of $T$. For $α>1$, the ratio $u_{\rm epi}/u_{\rm var}=1/ν_{\rm can}$ is therefore determined by the quotient state and prior; for rank-only use under a common calibration, $T$ contains the same coordinatewise ordinal information. The residual prior gauge is characterized rather than eliminated: $(γ_0,ρ_0)$ govern ordinal dependence, while $(ν_0,α_0)$ determine numerical calibration and the analytic ceiling of $1/ν_{\rm can}$.

发表机构

  • Rensselaer Polytechnic Institute(伦斯勒理工学院)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑