AI 中文总结
针对含分块缺失的垂直分布式数据,提出 ALB 算法,无需合并记录即可实现高维线性估计与推理,性能优于完整样本 Lasso,可近似集中式基准。
AI 中文摘要
在多机构研究中,不同参与方持有部分重叠个体集合的不同特征块,部分记录的响应也可能缺失。针对此类场景,我们提出分块缺失数据辅助学习算法(ALB),用于稀疏高维线性估计与坐标式推理,无需合并记录或依赖协调服务器。ALB 采用循环分块更新最小化正则化可用样本二次损失,每个循环通过样本级线性摘要通信 O(n) 个标量,与数据维度 p 无关,且迭代以几何速率收敛至集中式解。我们推导了分离统计误差与优化误差的估计速率。针对目标系数的推理,ALB 估计对应精度列,并采用考虑重叠样本计算的矩之间依赖性的样本级方差估计器。在稀疏性与重叠条件下,即使不存在完整样本,学生化估计器也能以√n 速率渐近服从标准正态分布。我们还研究了一次性扰动协变量与响应发布,通过用带噪版本替换未扰动的样本级量来减少直接披露。模拟实验与多模态阿尔茨海默病神经影像倡议数据分析表明,ALB 近似其集中式基准,且通过纳入部分观测记录优于完整样本 Lasso。
英文摘要
In multi-institutional studies, different parties hold distinct feature blocks for partially overlapping sets of individuals. Responses may also be missing for some records. In such settings, we propose Assisted Learning with Block-Missing Data (ALB) for sparse high-dimensional linear estimation and coordinatewise inference without pooling records or relying on a coordinating server. ALB minimizes a regularized available-case quadratic loss using cyclic block updates. Each cycle communicates $O(n)$ scalars through sample-level linear summaries, regardless of data dimension $p$, and the iterates converge geometrically to the centralized solution. We derive estimation rates that separate statistical and optimization errors. For inference on a target coefficient, ALB estimates the corresponding precision column and uses a sample-level variance estimator that accounts for dependence among moments computed from overlapping samples. Under sparsity and overlap conditions, the studentized estimator is asymptotically standard normal at the $\sqrt n$ rate, even when there are no complete cases. We also study one-time perturbed covariate and response releases that reduce direct disclosure by replacing unperturbed sample-level quantities with noisy versions. Simulations and an analysis of multimodal Alzheimer's Disease Neuroimaging Initiative data indicate that ALB approximates its centralized benchmark and improves upon complete-case Lasso by incorporating partially observed records.