发表机构
Friedrich-Alexander-Universität Erlangen-Nürnberg(埃尔朗根-纽伦堡大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
提出结合低秩初始化与随机森林细化的混合插补方法NuclearForest和SoftForest,在保持插补精度的同时,显著降低计算成本,实现高效稳健的表格数据插补。
AI 中文摘要
缺失数据是统计分析和机器学习中的一个基本挑战,因为插补方法的选择会显著影响下游推断。在本工作中,我们提出了两种混合插补方法,分别称为NuclearForest和SoftForest,它们将基于核范数的低秩初始化(分别使用奇异值阈值(SVT)和SoftImpute)与非迭代的随机森林细化相结合。对于基于SVT的组件,我们进一步引入了自适应步长规则,证明了自适应步长界限,并建立了相应零初始化迭代的收敛性。低秩初始化提供了结构化的热启动,捕捉数据中的全局协方差模式,而随后的随机森林步骤则恢复编码局部依赖的残差非线性信号。我们在来自不同应用领域的多样化数据集上进行了广泛的基准测试,将所提出的方法与七种已建立的插补方法在完全随机缺失(MCAR)、随机缺失(MAR)和非随机缺失(MNAR)机制下,在不同缺失率下进行了比较。我们的结果表明,NuclearForest和SoftForest在插补保真度上达到或超过了如MissForest等最先进的迭代方法,同时显著降低了计算成本。特别是,通过用单次细化步骤替代迭代循环,它们相对于MissForest实现了约5.81倍和9.52倍的加速。我们的方法有效利用了现实世界表格数据的低秩结构,并适应混合类型变量,为生物信息学、经济学及其他领域的数据插补提供了高效且稳健的解决方案。
英文摘要
Missing data are a fundamental challenge in statistical analysis and machine learning, as the choice of imputation method substantially impacts downstream inference. In this work, we propose two hybrid imputation methods called NuclearForest and SoftForest, which combine nuclear-norm-based low-rank initialization using Singular Value Thresholding (SVT) and SoftImpute, respectively, with a non-iterative Random Forest refinement. For the SVT-based component, we further introduce an adaptive step-size rule, prove adaptive step-size bounds, and establish convergence for the corresponding zero-initialized iteration. The low-rank initialization provides a structured warm start that captures the global covariance patterns in the data, while the subsequent Random Forest step recovers residual nonlinear signals encoding local dependencies. We conduct an extensive benchmark on diverse datasets from different application domains, comparing the proposed methods with seven established imputation methods under the Missing Completely at Random (MCAR), Missing at Random (MAR), and Missing Not at Random (MNAR) mechanisms across varying missingness rates. Our results demonstrate that NuclearForest and SoftForest match or exceed the imputation fidelity of state-of-the-art iterative methods such as MissForest, while significantly reducing computational cost. In particular, they achieve speedups of approximately 5.81 times and 9.52 times over MissForest by replacing iterative cycles with a single refinement step. Our approach effectively exploits the low-rank structure of real-world tabular data and accommodates mixed-type variables, providing an efficient and robust solution for data imputation in bioinformatics, economics, and beyond.
Comments35 pages, 8 figures