模式坍缩易于检测:神经采样器的无真值预检
Mode Collapse Is Cheap to Detect: A Ground-Truth-Free Pre-Flight Check for Neural Samplers
查看机构详情
- RIKEN iTHEMS(理化学研究所综合理论科学中心)
- RIKEN Center for Advanced Intelligence Project (AIP)(理化学研究所先进智能研究中心)
机构由 AI 辅助整理,请以论文原文为准。
浏览论文内容
中文总结 AI 辅助
本文提出一种仅利用目标能量及其梯度和Hessian的预检方法,以极低训练成本检测神经采样器是否遗漏目标质量,并在多种基准上验证其有效性与局限性。
中文摘要 AI 辅助
神经采样器针对未归一化目标 $\tilde\pi=e^{-E}$ 进行训练,且没有来自 $\pi$ 的样本,这使得实践者无法判断一次昂贵的训练运行是否静默地丢失了目标的一部分。常用诊断方法是从模型自身的样本中计算的,因此仅限于模型的支持域:我们展示了一个采样器,其自归一化有效样本量为 $0.99$,却遗漏了目标质量的 $87\\%$。我们认为,检测缺失质量在严格意义上比采样它更容易:检测每个遗漏盆地只需一个点加上局部曲率估计,而修正则需要重新训练采样器。我们将此转化为一种预检,消耗采样器自身训练预算的百分之几,且仅使用 $E$、$\nabla E$ 和 $\nabla^2 E$。在具有精确可计算真值的高斯混合、多井和旋转各向异性多井目标上,该检查以 $2.7\\%$ 的训练成本估计缺失质量至 $10^{-3}$ 以内,而调优的退火SMC参考需要 $70$--$280\\%$ 的训练成本才能做得更差。它同样适用于没有可处理密度的受控SDE采样器,其中ESS和ELBO根本无法形成。该估计器带有自诊断功能,在无真值情况下,在安全方向上保守:在 $60$ 个配置中,它通过了 $16$ 个,其中 $15$ 个准确度达到 $10^{-2}$ 或更好。我们明确说明这允许和不允许什么:该检查廉价地产生缺失质量的证据,有时产生搜索已稳定的证据,但它不能认证一次运行,其阈值是启发式的。然后,我们在真实物理景观LJ-13上绘制了该方法的边界,并报告了其失败之处及原因。
英文摘要
Neural samplers are trained against an unnormalised target $\tildeπ=e^{-E}$ with no samples from $π$, which leaves the practitioner with no way to tell whether an expensive training run has silently dropped part of the target. The diagnostics in common use are computed from the model's own draws and are therefore confined to the model's support: we exhibit a sampler whose self-normalised effective sample size is $0.99$ while it misses $87\%$ of the target mass. We argue that \emph{detecting} missing mass is a strictly easier problem than sampling it: detection needs one point per missed basin plus a local curvature estimate, whereas correction needs the sampler retrained. We turn this into a pre-flight check that consumes a few percent of the sampler's own training budget and uses only $E$, $\nabla E$ and $\nabla^2 E$. On Gaussian-mixture, Many-Well and rotated anisotropic Many-Well targets with exactly computable ground truth, the check estimates the missing mass to within $10^{-3}$ at $2.7\%$ of training cost, where a tuned annealed SMC reference needs $70$--$280\%$ of training cost to do worse. It also applies unchanged to a controlled-SDE sampler that has no tractable density, where ESS and the ELBO cannot be formed at all. The estimator carries a \emph{self-diagnostic} that, without ground truth, is conservative in the safe direction: across $60$ configurations it clears $16$, of which $15$ are accurate to $10^{-2}$ or better. We are explicit about what this does and does not license: the check cheaply produces evidence of missing mass, and sometimes evidence that the search has stabilised, but it cannot certify a run, and its thresholds are heuristic. We then map the boundary of the method on a real physical landscape, LJ-13, and report where it fails and why.