发表机构
Humboldt-Universität zu Berlin; Freie Universität Berlin; Zuse Institute Berlin; Technische Universität Berlin(柏林洪堡大学; 柏林自由大学; 柏林祖泽研究所; 柏林工业大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
LinearPFN是一种基于Transformer的摊销方法,用于线性模型(含交互项)的尖峰-平板变量选择,通过单次前向传播高效获取后验信息,在真实数据上优于经典基线。
AI 中文摘要
尖峰-平板回归是变量选择的一种标准贝叶斯表述:它返回候选效应活跃性的后验分布,而非单一选定的子集,因此每个候选效应都带有包含概率。其成本随候选效应数量呈指数增长,因此仅当预测变量数量较少时,后验才能被精确枚举。超出该范围后,后验必须近似计算,通常通过模型空间上的马尔可夫链蒙特卡罗方法进行,这需要对每个数据集进行全新运行,并且在固定的步数预算内可能无法收敛。我们提出LinearPFN,一种先验数据拟合的Transformer网络,用于摊销具有主效应和成对交互项的线性模型的尖峰-平板推断。该网络在从明确指定的先验中抽取的合成数据集上预训练一次,对新的数据集进行单次前向传播即可返回后验包含概率、后验均值系数和后验预测分布,无需针对每个数据集进行拟合。先验设计为共轭的,因此每个固定活跃效应集的后验具有封闭形式,并且在精确后验仍可通过枚举计算的地方,我们验证网络的输出。在已发表的社会科学数据集的真实预测变量矩阵上,结果来自先验以确认真实活跃集,LinearPFN在每次数据集的选择AUC和中位数概率模型规则下的F1均高于五个经典基线。当系数、交互项或噪声偏离先验时,这一优势依然保持。代码:此https URL。训练模型:此https URL。
英文摘要
Spike-and-slab regression is a standard Bayesian formulation of variable selection: it returns a posterior distribution over which candidate effects are active rather than a single selected subset, so that every candidate effect carries an inclusion probability. Its cost grows exponentially with the number of candidate effects, so the posterior can be enumerated exactly only when the number of predictors is small. Beyond that reach, the posterior has to be approximated, typically by Markov chain Monte Carlo over the model space, which requires a fresh run for every dataset and, within a fixed budget of steps, may fail to converge. We present LinearPFN, a prior-data fitted transformer network that amortizes spike-and-slab inference for linear models with main effects and pairwise interactions. The network is pretrained once on synthetic datasets, drawn from an explicitly specified prior, and a single forward pass over a new dataset returns posterior inclusion probabilities, posterior-mean coefficients and posterior predictive distributions with no per-dataset fitting. The prior is conjugate by design, so that the posterior for each fixed set of active effects has a closed form, and wherever the exact posterior is still computable by enumeration we verify the network's outputs against it. On real predictor matrices from published social-science datasets, with outcomes drawn from the prior so that the true active set is known, LinearPFN attains a higher per-dataset selection AUC and a higher F1 under the median probability model rule than five classical baselines. The lead holds when the coefficients, the interactions or the noise depart from the prior. Code: https://github.com/schiekiera/LinearPFN. Trained model: https://huggingface.co/schiekiera/LinearPFN.
Comments26 pages, 7 figures. Code: https://github.com/schiekiera/LinearPFN