arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.10398cs.LGcs.AI

ELVAE:用于不确定性感知生成的基于证据学习的变分自编码器

ELVAE: Evidential Learning-Based Variational Autoencoder for Uncertainty-Aware Generation

Ge Wang

首次发表
浏览论文内容

中文总结 AI 辅助

该研究提出基于证据学习的变分自编码器ELVAE,通过显式建模潜在位置不确定性,在MNIST生成实验中验证其可分层生成样本的语义可靠性,为不确定性感知生成提供实用控制变量。

中文摘要 AI 辅助

变分自编码器(VAE)从概率潜在表示中生成样本,但无法区分潜在位置的不确定性与围绕该位置的变异性。我们提出了ELVAE,这是一种基于证据学习的VAE,其中每个潜在坐标由依赖于输入的正态-逆伽马后验分布控制。该层次结构产生了显式的潜在位置不确定性,可在生成过程中使用,而非仅在推理后报告:低不确定性锚点支持更可靠的合成样本,而高不确定性锚点可被刻意用于压力测试。目标是精确的证据下界,且我们表明需要对整个层次结构进行直接正则化,因为仅边缘化的潜在规律无法识别不确定性分解。在使用冻结外部分类器的MNIST生成试点实验中,该不确定性清晰地分层了生成数字的语义可靠性。零位移对照实验显示,大部分效应反映了锚点可被重新生成的可靠性,而较小但显著的部分归因于不确定性缩放扰动本身。该效应仅在类内不确定性排序下成立,且其幅度随随机种子变化。这些发现支持学习到的潜在位置不确定性作为不确定性感知生成的实用控制变量,将锚点可靠性与扰动诱导的失效分离开来。

英文摘要

ELVAE places an input-dependent normal--inverse-gamma (NIG) hierarchy at each VAE latent coordinate, separating location uncertainty $u_{\mathrm{epi}}=β/[ν(α-1)]$ from conditional variability $u_{\mathrm{var}}=β/(α-1)$. The marginalized latent law, however, identifies only the three quotient coordinates $(γ,α,c)$ with $c=β(1+1/ν)$; reconstruction is blind to one $(ν,β)$ fiber direction. A companion theoretical analysis shows that the complete NIG prior and forward KL select a unique prior-relative canonical representative on each fiber, so canonical inverse allocation is not a fourth independent information channel. Empirically, trained inverse evidence $1/ν$ remains the strongest sensitivity-ranking score. At $τ_{\mathrm{epi}}=1$, the three-seed mean high/low-$u_{\mathrm{epi}}$ semantic-transition ratios are 1.98 on MNIST and 1.66 on Fashion-MNIST, falling to 1.33 and 1.16 under scale-matched controls. In the 20-draw MNIST component study with equal per-anchor perturbation energy, $1/ν$ gives high/low ratios 1.69 and 1.65 under the $u_{\mathrm{epi}}$ and $u_{\mathrm{var}}$ fields and 1.69 (95\% interval 1.45--1.97) under a geometry-free isotropic field, whereas $u_{\mathrm{var}}$ reverses the isotropic ordering to 0.76. Under the experimental prior, canonical $1/ν_{\mathrm{can}}$ is a strictly increasing transform of $T=c/[α(γ^2+2)]$ and is bounded above by $3+\sqrt{10}$. Thus ELVAE exposes a controllable sensitivity mechanism whose trained four-output realization is operationally informative, while the exact reconstruction-visible information remains three-dimensional and baseline image quality is a separate question.

发表机构

  • Rensselaer Polytechnic Institute(伦斯勒理工学院)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑