发表机构
Cornell University; Weill Cornell Medicine(康奈尔大学; 威尔康奈尔医学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对TMLE无法联邦化的问题,提出首个联邦TMLE算法FedTMLE-G与FedTMLE-L,通过梯度聚合或局部拟合交换实现分布式定向推断,并分析通信精度、隐私与收敛权衡。
AI 中文摘要
科学或运营决策背后的证据往往由医院、银行或登记机构持有,这些机构无法汇集个体观测数据。跨机构联邦学习将计算移至数据端,并交换商定的汇总信息。定向最大似然估计(TMLE)对灵活的初始拟合进行精炼,产生尊重模型并支持高效推断的插入式估计量。然而,TMLE本身仍然是一个完全集中式的过程。为填补这一空白,本文提出了首个联邦TMLE算法。我们通过两个互补框架,针对任意目标、损失函数和波动族,对定向过程本身进行联邦化。FedTMLE-G聚合局部梯度,并逐步复现集中式定向步骤。FedTMLE-L允许每个机构在单次交换拟合更新之前完成自身的波动拟合,以同步保真度换取局部自主性。对于梯度聚合,我们开发了一种有限精度协议,该协议传输变化量而非原始值,并在交换次数和比特数的显式界限内保证定向精度。对已接受更新的描述长度分析表明,这种有限通信使数值定向误差相对于采样不确定性可忽略不计。因此,计算估计量的成本与选择估计量的复杂性是不同的。我们的分析还表明,将数据保留在本地本身并非TMLE的隐私保证,因为全记录重建的不稳定性并不妨碍恢复特定敏感属性。对于个性化版本的局部平均,机构保留各自的估计,并在局部定向完成后退出。一个非凸收敛界将因平均而损失的改进归因于局部拟合之间的分歧,并揭示了机构影响力平等与小数据孤岛采样变异性之间的权衡。
英文摘要
The evidence behind a scientific or operational decision is often held by hospitals, banks, or registries that cannot pool individual observations. Cross-silo federated learning moves computation to the data and exchanges agreed summaries. Targeted maximum likelihood estimation (TMLE) refines a flexible initial fit, yielding plug-in estimators that respect the model and support efficient inference. TMLE itself, however, has remained a fully centralized procedure. To fill this gap, our paper introduces the first federated TMLE algorithm. We federate targeting itself, for an arbitrary target, loss, and fluctuation family, through two complementary frameworks. FedTMLE-G aggregates local gradients and reproduces centralized targeting step for step. FedTMLE-L lets each institution complete its own fluctuation fit before a single exchange of fitted updates, trading synchronized fidelity for local autonomy. For gradient aggregation, we develop a finite-precision protocol that transmits changes rather than values and certifies targeting accuracy within explicit bounds on exchanges and bits. A description-length analysis of the accepted updates then shows that this finite communication leaves numerical targeting error negligible against sampling uncertainty. The cost of computing an estimator is thus distinct from the complexity of selecting it. Our analysis also indicates that keeping data local is not itself a privacy guarantee of TMLE, since instability of full-record reconstruction need not prevent recovery of a specified sensitive attribute. For a personalized version of local averaging, institutions retain their own estimates and leave once local targeting is complete. A nonconvex convergence bound charges the improvement forfeited through averaging to disagreement among local fits and exposes a tradeoff between equal institutional influence and the sampling variability of small silos.