arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

用于流式联邦学习的自适应数据接纳与保留

Adaptive Data Admission and Retention for Streaming Federated Learning

Zhuoyi Zhao, Ben Liang

arXiv 2607.23987首次发表:更新:

发表机构

University of Toronto(多伦多大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究客户端内存有限的流式联邦学习,提出联合服务器端接纳和客户端内存管理框架,开发ACDPP策略,结合K步保留规则与在线接纳规则等,通过论证使其与预言机基准联系,实验表明该策略在满足约束时接近基准。

AI 中文摘要

我们研究了客户端内存有限的流式联邦学习,其中新生成的训练数据会产生随时间变化的采样成本,且必须随时间选择性地接纳和保留。我们考虑一个联合的服务器端接纳和客户端内存管理框架,目标是在采样成本预算和缓冲区约束下最小化累积超额总体风险。我们首先通过有效样本大小的表征得出一个学习误差界,该界明确捕捉了瞬时训练样本大小、不同样本增长和重用不平衡的影响。通过从该界获得的替代惩罚,我们开发了一种主动约束漂移加惩罚(ACDPP)策略,它将结构化的客户端K步保留规则与服务器端在线接纳规则及时变矩形接纳区域相结合。我们还通过辅助常数接纳策略给出了一系列比较论证,将ACDPP学习界与无成本预言机基准联系起来。这在次线性遗憾和采样成本违反方面给出了明确保证,而缓冲区占用违反则通过离线选择保留期来控制。在多个数据集上的实验表明,所提出的策略在满足采样成本和缓冲区约束的同时,仍接近预言机基准。

英文摘要

We study streaming federated learning with limited client memory, where newly generated training data incur time-varying sampling costs and must be selectively admitted and retained over time. We consider a joint server-side admission and client-side memory-management framework with the objective of minimizing the cumulative excess population risk under a sampling-cost budget and buffer constraints. We first derive a learning-error bound that explicitly captures the effects of instantaneous training sample size, distinct-sample growth, and reuse imbalance through a characterization of the effective sample size. Through a surrogate penalty obtained from this bound, we develop an Active-Constraint Drift-Plus-Penalty (ACDPP) policy that combines a structured client-side $K$-step retention rule with a server-side online admission rule and a time-varying rectangular admission region. We further present a sequence of comparison arguments, via an auxiliary constant-admission policy, that connects the ACDPP learning bound to a costless oracle benchmark. This yields explicit guarantees in terms of sublinear regret and sampling-cost violation, while the buffer-occupancy violation is controlled through offline selection of the retention horizon. Experiments on multiple datasets demonstrate that the proposed policy remains close to the oracle benchmark while satisfying the sampling-cost and buffer constraints.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑