学习三角传输映射的结构
Learning the Structure of Triangular Transport Maps
- University of Bergen(卑尔根大学)
- Equinor(挪威国家石油公司)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本文提出自结构化传输映射(SSTM),联合学习三角映射的排序、稀疏性与密度,使用SoftSort和L0门控,在合成及真实数据上优于先估计结构的方法,并匹配或超越自回归流。
AI中文摘要:
三角传输映射为基于采样的概率建模提供了一种灵活的方法,包括密度估计、生成建模和贝叶斯推断。它们通过单调三角映射将未知的目标分布变换为更简单的参考分布。映射结构由变量排序和稀疏模式定义,这两者共同编码了一个有向无环图。映射的质量在很大程度上依赖于该结构,然而找到好的结构在计算上代价高昂,因为每个候选结构通常需要拟合不同的映射。因此,一个核心挑战是在维度增长时,联合学习密度和结构,同时保持计算的可管理性。我们引入了自结构化传输映射(SSTM),它联合学习映射、排序和稀疏性。我们使用SoftSort来学习变量排序,使用$L_0$门控来学习稀疏性,同时保持三角结构。为了使映射具有可扩展性,我们使用单调的BatchEnsemble,通过秩一适配器在所有映射组件之间共享一个权重矩阵。在合成数据和真实数据上,联合学习结构和映射比先估计结构再学习映射能获得更好的密度估计。当结构可从密度中识别时,SSTM在密度性能上匹配使用真实结构拟合的映射,并优于自回归流。在大型数据集上,SSTM与自回归流具有竞争力。
英文摘要:
Triangular transport maps provide a flexible approach to sampling-based probabilistic modeling, including density estimation, generative modeling, and Bayesian inference. They transform an unknown target distribution into a simpler reference through a monotone triangular map. The map structure is defined by a variable ordering and sparsity pattern, which together encode a directed acyclic graph. Map quality can depend strongly on this structure, yet finding a good structure is computationally expensive because each candidate generally requires fitting a different map. A central challenge is therefore to learn density and structure jointly, while keeping computation manageable as dimension grows. We introduce Self-Structuring Transport Maps (SSTM), which learn the map, ordering, and sparsity jointly. We use SoftSort to learn the variable ordering and $L_0$ gates to learn the sparsity, while preserving a triangular structure. To keep the map scalable, we use a monotone BatchEnsemble that shares one weight matrix across all map components through rank-one adapters. Across synthetic and real data, jointly learning the structure and map gives better density estimates than estimating the structure first. When the structure is identifiable from the density, SSTM matches the density performance of a map fitted with the true structure and outperforms autoregressive flows. On large datasets, SSTM is competitive with autoregressive flows.