压缩活性子空间用于可扩展贝叶斯推断
Compressed Active Subspaces for Scalable Bayesian Inference
- Brookhaven National Laboratory(布鲁克海文国家实验室)
- Texas A&M University(德克萨斯A&M大学)
- Argonne National Laboratory(阿贡国家实验室)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
针对活性子空间方法在大模型中存储梯度内存过大的问题,提出压缩活性子空间(CAS),通过结构化等距嵌入降维后构建子空间,大幅降低内存需求,实现大规模神经网络的贝叶斯推断并保持性能与不确定性估计。
AI中文摘要:
活性子空间方法通过识别模型输出影响最大的参数方向并沿这些方向进行推断,为高维模型中的预测不确定性量化提供了框架。然而,活性子空间的构建需要存储大量全维模型梯度,随着模型规模的增大,这一需求变得难以承受。我们通过提出压缩活性子空间(CAS)来解决这一局限,这是一种可扩展的方法,首先使用结构化等距嵌入将模型参数映射到压缩空间,然后在此降维参数化内构建活性子空间。我们的方法大幅减少了活性子空间构建所需的内存,使得标准活性子空间方法不可行的大规模模型也能进行贝叶斯推断。我们在规模递增的神经网络上展示了CAS的可扩展性,同时保持了预测性能和稳健的不确定性估计。
英文摘要:
Active subspace methods provide a framework for quantifying predictive uncertainty in high-dimensional models by identifying and performing inference along parameter directions that have the greatest influence on the model output. However, the construction of active subspaces requires storing many full-dimensional model gradients, which becomes prohibitive as model size increases. We address this limitation by proposing Compressed Active Subspaces (CAS), a scalable approach that first maps the model parameters to a compressed space using a structured isometric embedding and then constructs the active subspace within this reduced parameterization. Our approach substantially reduces the memory required for active subspace construction and enables Bayesian inference for large models where standard active subspace methods become impractical. We demonstrate the scalability of CAS on neural networks of increasing size while maintaining predictive performance and robust uncertainty estimates.