发表机构
Purdue University; IBM Research(普渡大学; IBM研究院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对现代选择性状态空间模型(如Mamba2)在联邦学习中的架构无关问题,推导架构感知的梯度和收敛界,并通过数值验证及九种算法在六个文本域上的实验,揭示稳定性、离散化与投影范数对联邦优化的影响。
AI 中文摘要
现代状态空间模型(SSMs),如Mamba2,通过将线性时间序列建模与循环状态空间动态相结合,为Transformer提供了一种引人注目的替代方案。然而,SSMs在分布式学习环境中的行为仍鲜为人知。特别是,现有的标准联邦学习方法在很大程度上是架构无关的,并未考虑现代选择性SSMs所具有的稳定性、选择性以及状态空间参数化特征。为解决这一问题,我们推导了针对单层和多层选择性SSMs的架构感知的梯度和光滑性界,以及FedAvg和FedProx的收敛界,刻画了循环稳定性、输入依赖的离散化以及状态投影范数如何影响联邦优化。随后,我们在由教师SSM生成的序列上数值验证了单层界,使用遵循所分析循环的学习器。我们利用这一分析来形成关于本地训练和客户端异质性影响的预期,并通过在六个文本领域上比较九种联邦学习算法在Mamba2语言建模中的表现来检验这些预期。这些实验说明了SSM特有的界如何为解释实际联邦学习算法的行为提供基础。
英文摘要
Modern state space models (SSMs), such as Mamba2, provide a compelling alternative to transformers by combining linear-time sequence modeling with recurrent state-space dynamics. However, the behavior of SSMs in distributed learning settings remains poorly understood. In particular, the existing standard federated learning methods are largely architecture-agnostic, and do not account for the stability, selectivity, and state-space parameterization that characterize modern selective SSMs. To address this, we derive architecture-aware gradient and smoothness bounds for single- and multi-layer selective SSMs, and convergence bounds for FedAvg and FedProx, characterizing how recurrent stability, input-dependent discretization, and state projection norms affect federated optimization. We then numerically validate the single-layer bounds on sequences generated by a teacher SSM, using a learner that follows the analyzed recurrence. We use this analysis to formulate expectations about the effects of local training and client heterogeneity, and examine these expectations by comparing nine federated learning algorithms on Mamba2 language modeling across six text domains. These experiments illustrate how SSM-specific bounds can provide a basis for interpreting the behavior of practical federated learning algorithms.