训练中去重:联邦学习中隐私保护跨客户端去重的弹性范式
Deduplication-while-Training: A Resilient Paradigm for Privacy-Preserving Cross-Client Deduplication in Federated Learning
浏览论文内容
中文总结 AI 辅助
针对联邦学习训练语料跨客户端重复数据导致的效率与隐私问题,提出训练中去重范式及系统DwT-FL,通过并发状态认领与热冷双队列调度实现去重与训练并行,显著降低故障恢复和动态加入开销。
中文摘要 AI 辅助
大型语言模型训练语料库中的跨客户端重复数据会降低联邦学习(FL)的效率,同时加剧模型记忆化和隐私风险。隐私保护的跨客户端去重通过消除重复训练数据有效缓解了这一问题。然而,现有方案均遵循“训练前去重”范式。这种串行耦合范式带来了高昂的容错成本,且不支持动态客户端加入。为此,我们提出了一种未被探索的范式,称为“训练中去重(DwT)”,该范式支持去重与训练并发执行。DwT将跨客户端去重从一次性的、全局同步的预处理操作转变为一种具有状态管理、并发认领和故障恢复的持续在线服务。通过实现状态同步和任务接管,它最大限度地减少了客户端断连对整体训练进度的影响,同时支持客户端的动态加入。我们设计了DwT-FL,一个支持DwT的隐私保护去重系统。通过设计并发状态认领机制和热-冷双队列调度策略,DwT-FL实现了安全去重与模型训练的并行执行,同时有效处理客户端断连和动态加入。实验评估表明,与最先进方案相比,DwT-FL在故障恢复和动态加入的时间开销上分别显著降低了高达93.04%和94.18%。这为动态且不稳定的FL环境提供了一种高效且弹性的并发去重方案。
英文摘要
Cross-client duplicate data in large language model training corpora degrades the efficiency of federated learning (FL) while exacerbating model memorization and privacy risks. Privacy-preserving cross-client deduplication effectively mitigates this issue by eliminating duplicate training data. However, existing schemes all follow a "Deduplication-before-Training" paradigm. This serially coupled paradigm incurs high fault-tolerance costs and lacks support for dynamic client joining. To this end, we propose an unexplored paradigm called "Deduplication-while-Training (DwT)", which enables concurrent deduplication and training. DwT transforms cross-client deduplication from a one-time, globally synchronous preprocessing operation into a continuous online service with state management, concurrent claiming, and failure recovery. By enabling state synchronization and task takeover, it minimizes the impact of client disconnections on the overall training progress while supporting the dynamic joining of clients. We design DwT-FL, a privacy-preserving deduplication system, to support DwT. By designing a concurrent state-claim mechanism and a hot-cold dual-queue scheduling strategy, DwT-FL enables the parallel execution of secure deduplication and model training, while effectively handling client disconnections and dynamic joins. Experimental evaluations demonstrate that, compared to the state-of-the-art scheme, DwT-FL significantly reduces the time overhead of failure recovery and dynamic joining by up to 93.04% and 94.18%, respectively. This provides an efficient and elastic concurrent deduplication scheme for dynamic and unstable FL environments.
发表机构
- College of Cryptology and Cyber Science, Nankai University(南开大学密码学与网络科学学院)
- School of Mathematical Science, Nankai University(南开大学数学科学学院)
机构由 AI 辅助整理,请以论文原文为准。