arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.19585stat.MEstat.AP

基于Vine Copula插补的数据融合估计负收入分布

Estimating Negative Income Distributions via Data Fusion with Vine Copula-based Imputation

Sithara Wijekoon, Helen Thompson, Summer Wang, Gentry White

AI总结:

本研究提出基于C-Vine和D-Vine Copula条件抽样的插补数据融合框架,经模拟与真实数据验证,其在保留插补分布和依赖结构上优于现有方法,可用于估计负收入分布。

AI中文摘要:

基于插补的数据融合通过利用一个数据集的信息插补另一个数据集的缺失变量来合并数据集,其有效性依赖于准确保留插补分布和跨数据源的潜在多元依赖结构。然而,传统基于插补的数据融合方法在保留依赖结构方面常存在困难,尤其是当边际分布不同或存在复杂非线性多元依赖时,依赖结构保留不足会导致联合分布失真和融合数据的推断偏差。为解决这一挑战,本研究提出一种新颖的基于插补的数据融合框架,利用C-Vine和D-Vine Copula的条件抽样作为插补机制,该方法可灵活建模成对及高阶依赖,同时适配异质边际分布,从而生成更准确维持多元数据分布的插补值。所提方法通过调查数据的模拟研究和真实数据应用验证,即利用行政税务记录的信息估计调查数据中的负收入分布,结果表明该方法在保留跨数据集的插补分布和依赖结构方面优于现有方法。

英文摘要:

Imputation-based data fusion combines datasets by imputing missing variables in one dataset using information from the other. The validity of imputation-based data fusion relies on accurately preserving both the imputed distributions and the underlying multivariate dependence structure across data sources. However, traditional imputation-based data fusion methods often struggle to preserve the dependence structure, particularly when marginal distributions differ or when complex, nonlinear multivariate dependencies exist. Inadequate preservation of these dependencies can result in distorted joint distributions and biased inference in the fused data. To address this challenge, this study introduces a novel imputation-based data fusion framework that utilises conditional sampling from C-vine and D-vine copulas as the imputation mechanism. This approach flexibly models pairwise and higher-order dependencies while accommodating heterogeneous marginal distributions, thereby generating imputations that more accurately maintain the distribution of the multivariate data. The proposed method is validated through a simulation study using survey data and a real-data application that estimates the negative income distribution in survey data using information from administrative tax records. Results indicate that the proposed method outperforms existing approaches in preserving both the imputed distributions and the dependence structure across datasets.

↑