利用辅助监督缓解摊销贝叶斯推断中的表示差距
Mitigating Representation Gaps in Amortized Bayesian Inference with Auxiliary Supervision
查看机构详情
- TU Dortmund University(多特蒙德工业大学)
- Rensselaer Polytechnic Institute(伦斯勒理工学院)
机构由 AI 辅助整理,请以论文原文为准。
浏览论文内容
中文总结 AI 辅助
本文提出利用辅助监督损失改善摊销贝叶斯推断的训练动态,缓解表示差距,加快收敛并提升性能,并在多个现实问题中验证了有效性。
中文摘要 AI 辅助
将贝叶斯推断视为针对摊销后验的神经网络优化问题具有吸引力,因为它扩展到原本难以处理的统计模型,并在支付训练成本后为新的数据集提供近乎即时的推断。尽管理论保证在理想收敛下的忠实性,实际摊销推断仍需要迭代架构和优化选择,并最终在有限的模拟、计算和时间预算下“满意”。即使是最佳解决方案也可能保留可避免的表示差距,这些差距通常需要特定问题的修复。在这里,我们提出了一种通用替代方案,通过应用于内部表示的辅助指导损失来改善训练动态。具体来说,我们展示了这种指导如何在训练数据充足时导致更快的收敛,在数据稀缺时导致更好的性能。我们将表示差距形式化为在信息瓶颈处陷入局部最优,该瓶颈位于负责特征学习的网络部分与负责条件分布学习的部分之间,并提供了一种通用诊断方法来区分摘要失败与推断失败。最后,我们证明辅助监督在一系列具有挑战性的现实世界推断问题上提高了收敛速度和准确性。
英文摘要
Casting Bayesian inference as a neural network optimization problem targeting an amortized posterior is attractive, as it extends to otherwise intractable statistical models and offers near instantaneous inference for new datasets after prepaying the training cost. Although theory guarantees faithfulness under ideal convergence, practical amortized inference still requires iterating over architectures and optimization choices and ultimately ``satisficing'' under finite simulation, compute, and time budgets. Even the best-performing solution may thus retain avoidable representation gaps that typically require problem-specific fixes. Here, we propose a generic alternative which improves training dynamics with auxiliary guidance losses applied to internal representations. Specifically, we show how such guidance leads to faster convergence when training data is abundant and to better performance when it is scarce. We formalize representation gaps as getting stuck in a local optimum at the information bottleneck between the parts of the network tasked with feature learning and those tasked with conditional distribution learning, and offer a generic diagnostic to separate summary failures from inference failures. Finally, we demonstrate that auxiliary supervision improves convergence speed and accuracy on a range of challenging real-world inference problems.