arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.05025cs.LGstat.ML

CIFAR-10上的规范联合能量基模型:预测-校正采样器与SGLD采样器的失效模式及实际不可区分性

Comparing SGLD and a fixed-noise Predictor-Corrector adaptation in canonical Joint Energy-Based Models on CIFAR-10

Dmytro Knopov

首次发表
浏览论文内容

中文总结 AI 辅助

本研究复现CIFAR-10上的规范联合能量基模型,发现预测-校正采样器与SGLD采样器无方法级优势差异,同时记录到两类失效模式。

中文摘要 AI 辅助

联合能量基模型(JEM)将分类与生成统一在单个网络中,支持分布外(OOD)检测。规范JEM训练依赖随机梯度朗之万动力学(SGLD);理论上有动机的替代方案预测-校正(PC)采样器此前未在规范模型上进行过系统复现测试。我们在WideResNet-28-10上复现规范JEM,未使用归一化层,开展两次独立运行测试,探究PC在无退火噪声调度下是否保留理论优势,涉及三项协议:PC在约130个训练周期全程替换SGLD、冷启动生成(FID)、精调式多OOD检测(AUROC)。复现结果为测试准确率92.88%,缓冲FID为44.46(规范值为92.9%和38.40)。我们记录到两种失效模式:通过规范异常值缓冲机制产生的灾难性训练后期发散(SGLD两次运行及PC两次运行均出现相同特征),以及依赖运行的SVHN OOD判别动态。在所有协议中,均未观察到PC相较于SGLD的方法级优势:推理时,所有10个检查点-OOD对的绝对AUROC差值低于0.007,FID差值低于0.5;训练协议中,逐图像分层bootstrap给出宏平均AUROC差值的95%置信区间包含0,而每种方法两次运行的种子级等价检验无法确立形式等价性。数据与等价性及小方向效应均一致。这种实际不可区分性在理论上是可预期的:固定噪声下,PC预测步会按构造退化,因此其保证无法迁移到规范JEM。

英文摘要

Joint Energy-Based Models (JEM) unify classification and generation within a single network and support out-of-distribution (OOD) detection. Canonical JEM training relies on stochastic gradient Langevin dynamics (SGLD); a theoretically motivated alternative, the Predictor-Corrector (PC) sampler, has not previously undergone a systematic replication test on the canonical model. We reproduce canonical JEM on WideResNet-28-10 without normalisation layers on two independent runs and test a fixed-noise PC adaptation - with the degenerate annealed-noise predictor replaced by a deterministic gradient step - across three protocols: the adapted sampler replacing SGLD throughout the full training trajectories (115-132 epochs); cold-start generation (FID); and refinement-style multi-OOD detection (AUROC). The reconstruction reaches 92.88% test accuracy and buffer-FID 44.46 (canonical: 92.9% and 38.40). We document two failure modes: catastrophic late-training divergence with the signature of the canonical outlier-buffer mechanism (all four runs), and run-dependent SVHN OOD-discrimination dynamics. No consistent method-level advantage of the adaptation over SGLD is observed on any protocol: refinement AUROC differences stay below 0.007 across ten checkpoint-OOD pairs; seeded cold-start generation favours SGLD by about five FID points; on the training protocol a hierarchical seed-by-image bootstrap gives a 95% confidence interval on the macro-averaged AUROC difference that contains zero, while a seed-level equivalence test with two runs per method cannot establish formal equivalence. The training-protocol data are consistent both with equivalence and with a small directional effect. This outcome is consistent with theory: the guarantees of the annealed-noise PC framework do not transfer to the constant-noise regime of canonical JEM.

发表机构

  • National University of Kyiv-Mohyla Academy(基辅莫希拉国立大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑