arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.06946cs.DCcs.LG

流学习:无令牌的分区公平闲聊学习

Stream Learning: Partition-Fair Gossip Learning Without Tokens

  • Swapcard(Swapcard公司)
  • Sorbonne Université(索邦大学)
  • CNRS(法国国家科学研究中心)
  • Institut Universitaire de France(法国大学研究院)

机构由 AI 辅助整理,请以论文原文为准。

Fabien Mathieu, Alexandre Pham, Maria Gradinariu Potop-Butucaru, S{é}bastien Tixeuil

AI总结:

该研究提出流学习方法,其核心协议Ri无需令牌与元数据,在无故障场景下性能与PTGL相当,在节点崩溃的异构场景下表现更优。

AI中文摘要:

在闲聊学习中,节点网络无需中央协调器,通过反复交换本地模型的部分内容来协同训练共享模型。Hegedüs等人提出的最先进协议——分区令牌闲聊学习(PTGL),将权重矩阵划分为S个固定分区,并采用基于令牌的公平机制,结合每邻居元数据交换来传播这些分区。我们借鉴对等实时流媒体的类比重新审视分区调度,其中模型分区如同视频块,分区年龄如同块稀缺性。该类比衍生出两阶段选择策略的设计空间(先选分区或先选邻居),我们从中实例化出十种具体协议,统称为流学习。我们的主要发现是,其中最简单的协议——将本地训练最少的分区传输给均匀随机邻居(Ri),在无故障工作负载下与PTGL性能相当,且无需令牌计数器或元数据交换。在最佳性能节点遭受30%永久崩溃的对抗场景下,Ri在所有测试的完全图配置中与PTGL性能相当或更优,在最异构场景(Dirichlet β=0.1)下,HAR数据集上的差距达5.53%,MNIST数据集上达5.41%。实验表明,由分区年龄的单一局部规则所体现的分区公平性是差距产生的原因;基于令牌的速率控制和效用最大化并未优于该规则,且在异构场景下表现更差。

英文摘要:

In gossip learning, a network of nodes trains a shared model collaboratively, without a central coordinator, by repeatedly exchanging parts of their local models. The state-of-the-art protocol, Partitioned Token Gossip Learning (PTGL) of Heged{ü}s et al., splits the weight matrix into S fixed partitions and disseminates them using a token-based fairness mechanism coupled with per-neighbor metadata exchange. We revisit partition scheduling by analogy with peer-to-peer live streaming, where model partitions act as video chunks and partition age acts as chunk scarcity. The analogy yields a design space of two-stage selection strategies (partition first, or neighbor first), from which we instantiate ten concrete protocols collectively called Stream Learning. Our main finding is that the simplest of these protocols, which transmits the locally least-trained partition to a uniformly random neighbor (Ri), matches PTGL on fault-free workloads while requiring neither token counters nor metadata exchange. Under an adversarial 30% permanent crash of the best-performing nodes, Ri matches or outperforms PTGL across all complete-graph configurations tested, with the gap reaching 5.53% on HAR and 5.41% on MNIST in the most heterogeneous regime (Dirichlet $β$ = 0.1). In our experiments, partition fairness, captured by a single local rule on partition age, accounts for the gap; token-based rate control and utility maximization do not improve over this rule and, under heterogeneity, sit below it.

补充信息

↑