发表机构
Johns Hopkins University; The University of Texas at Austin; Yale University(约翰斯·霍普金斯大学; 德克萨斯大学奥斯汀分校; 耶鲁大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
Marformer是一种Transformer模型,可直接预测缺失变量的条件边际分布,在合成与真实缺失数据场景中,其性能优于或相当于经典方法,且速度快于生成式基线。
AI 中文摘要
实际决策是在信息不完整的情况下做出的。如果我们仅观察到所需的部分随机变量,就需要对其余变量进行预测。缺失变量上的条件边际分布(conditional marginals)是计算贝叶斯风险和信息价值(Value of Information, VOI)的关键要素,信息价值是指在决策前获取额外观测值的预期收益。我们提出了Marformer,这是一种Transformer模型,其训练目标是在给定任意一组观测值的情况下直接预测条件边际分布。与BERT(通过上下文预测缺失单词的模型)类似,Marformer为每个分布$p(X_i)$构建隐向量表示,并通过对其他分布$p(X_j)$的注意力机制迭代优化该表示。与生成式方法不同,Marformer无需对完整联合分布进行建模,不需要数据生成过程的领域知识,且所有预测仅需一次前向传播。我们在三个存在缺失数据的合成领域(贝叶斯网络、离散化多元高斯分布、结构化标注数据)进行了评估。Marformer的性能可与经典缺失数据方法相当甚至更优,即便这些方法拥有合成数据生成的真实模型族和先验知识。我们还在真实标注数据集上进行了评估,结果显示在最大训练规模下,Marformer的性能优于所有评估的基线方法。在两种场景中,Marformer的运行速度都显著快于评估的生成式基线方法。
英文摘要
Real decisions are made under incomplete information. If we observe only some of the random variables we need, we can predict the others. The \textbf{conditional marginals} over the missing variables are the key ingredient for computing Bayes risk and Value of Information (VOI), the expected gain from acquiring one more observation before deciding. We present the Marformer, a Transformer trained to directly predict conditional marginals given any set of observed values. Like BERT, which is trained to predict missing words from context, the Marformer constructs a hidden-vector representation for each distribution $p(X_i)$ and iteratively refines it through attention to other distributions $p(X_j)$. Unlike generative approaches, the Marformer does not model the full joint distribution, requires no domain knowledge of the data-generating process, and makes all predictions in a single forward pass. We evaluate across three synthetic domains with missing data---Bayesian networks, discretized multivariate Gaussians, and structured annotation data. The Marformer can match or outperform classical missing-data methods, even when those methods are given the true model family and prior that generated the synthetic data. We also evaluate on a real annotation dataset, where the Marformer outperforms the evaluated baselines at the largest training size. In both cases, the Marformer is substantially faster than the evaluated generative baselines.
Comments45 pages, 19 figures; presented in part in a COLM 2026 keynote