AI 中文总结
本文针对用大规模辅助样本扩充极小目标样本的问题,研究IPW与FL两种方法的效率特性,证明FL法可实现全效率提升,还探讨其在指数族、非参数框架及神经网络训练中的应用。
AI 中文摘要
本文研究用大规模辅助样本扩充极小目标样本的问题。利用Tukey分解,存在两种常用方法:逆概率加权(IPW)法和全似然(FL)法。我们表明,IPW法受限于目标样本规模小的问题,而FL法可按大规模辅助样本的速率估计部分模型参数,这种现象我们称为全效率提升。我们研究指数族及指数族混合模型下全效率提升的理论,还研究非参数框架下IPW法的效率提升情况,说明其如何达到目标样本规模对应的参数速率。作为补充说明,我们还讨论如何用FL法同时针对目标分布和优势比模型训练神经网络模型。
英文摘要
In this paper, we study the problem of augmenting a tiny target sample with a massive auxiliary sample. Utilizing Tukey's factorization, there are two popular approaches: the inverse probability weight (IPW) and the full-likelihood (FL) methods. We show that the IPW approach suffers from the limited target sample problem while the FL method may estimate some model parameters at the rate of the massive auxiliary sample size, a phenomenon we call full efficiency gain. We study the theory behind the full efficiency gain for exponential families and mixtures of exponential families. We also study the efficiency gain for the IPW method under a nonparametric procedure and show how it can achieve a parametric rate of the target sample size. As a side note, we also discuss how one may use FL to train neural network models simultaneously for both the target distribution and the odds model.
CommentsMain paper: 31 pages. 6 tables and 7 figures