发表机构
Meijo University(名城大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究提出观察块多预测,使深度玻尔兹曼机在统计数据融合中可用,并发现跨结果块条件化是性能提升的关键,且不随数据规模衰减。
AI 中文摘要
统计数据融合结合两个面板,它们共享一个协变量块,但观察不相交的结果块,在其传统形式中,没有一行同时观察两个结果。这排除了人们更希望用来训练深度玻尔兹曼机的判别性标准,因为多预测训练需要其保留的任何内容的真实值。我们提出了观察块多预测,它将多预测目标限制为从每行实际观察到的目标中抽取。它对于任何缺失模式都有良好定义,并且在行完整时简化为原始标准。拥有在该设置中存活的判别性标准使我们能够通过将联合模型的贡献分为表示部分和推理部分来询问是否根本需要联合模型。在两个消费者面板上,在覆盖35个单元格和875次运行的样本量和协变量宽度的网格上,微调后的DBM在每个单元格中都是十五种方法中最好的;但几乎所有这些优势都不来自生成式预训练,后者仅限于一个数据集上的最小样本量,在另一个数据集上则不存在。它来自于在预测一个结果块时以另一个结果块为条件。这一项相当于+0.19和+0.07个百分点,在所有35个单元格中均为正,并且与我们测量的其他贡献不同,它既不会随着面板的增长而衰减,也不需要第二个隐藏层,也不需要更多推理。置换一个结果块以破坏其与另一个结果块的关联会完全消除收益,这正是该解释所预测的。边际很小。但一个不会衰减的小效应与一个会衰减的效应是不同的对象,因为它基于的证据是,没有任何将协变量映射到结果的模型能够接受。
英文摘要
Statistical data fusion combines two panels that share a block of covariates but observe disjoint outcome blocks, and in its traditional form no row observes both outcomes at once. That rules out the discriminative criterion one would rather train a Deep Boltzmann Machine with, since multi-prediction training needs ground truth for whatever it holds out. We propose observed-block multi-prediction, which restricts the multi-prediction objective to targets drawn from what each row actually observes. It is well defined for any missingness pattern and reduces to the original criterion when rows are complete. Having a discriminative criterion that survives the setting lets us ask whether the joint model is needed at all, by separating what it contributes into a representation part and an inference part. On two datasets of different kinds, a consumer purchase panel and public-domain census microdata, over grids in sample size and covariate width spanning 40 cells and 200 runs per method, almost none of the fine-tuned DBM's advantage comes from generative pre-training, which is confined to the smallest sample size on one dataset and absent on the other. It comes from conditioning on one outcome block when predicting the other. This term amounts to +0.19 and +0.36 percentage points, is positive in all 40 cells, never decays as the panels grow (it is flat on one dataset and grows on the other), and requires neither a second hidden layer nor more inference. Against baselines tuned on validation and given the same conditioning, the fine-tuned DBM is the best method in 37 of the 40 cells. The imputers that can also condition on the other outcome block mostly lose accuracy when they do, whereas the DBM gains in every cell; since fusion data cannot validate that choice, this is the property that matters.