arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.06838cs.DC

StateFlow:用于长上下文建模的序列流水线并行方法,基于线性循环

StateFlow: Sequence Pipeline Parallelism for Long-Context Modeling with Linear Recurrence

Wenxuan Zhao, Yingfa Chen, Xu Han, Wenjing Han, Tianbo Huang, Zhiyu Li, Ao Sun, Jingheng Xu, Lin Gan, Guangwen Yang

首次发表
浏览论文内容

中文总结 AI 辅助

针对长上下文训练并行化难题,提出StateFlow序列流水线并行系统,通过分块调度、非均匀分块等优化,在32B参数、256K上下文模型上实现吞吐量与内存性能提升。

中文摘要 AI 辅助

长上下文训练对大型语言模型愈发重要,线性注意力和状态空间模型已成为提升长上下文效率的热门技术。然而,对循环模型及混合模型进行长序列训练的高效并行化仍具挑战。本文提出StateFlow,一种面向线性循环模型的序列流水线并行系统。StateFlow将每个序列划分为多个块,调度其执行的同时在块间传递边界状态与梯度,以此缩短激活生命周期、提升训练吞吐量。StateFlow进一步采用基于性能分析的非均匀分块策略,平衡混合模型中循环与softmax注意力的计算负载,并将并行度有限的状态转换与周围计算重叠执行。将StateFlow应用于参数规模达320亿、上下文长度达256K的模型,相较于传统流水线并行,可实现最高2.22倍的吞吐量提升与2.45倍的内存降低,支持原本不可行的配置。

英文摘要

Long-context training is increasingly important for large language models, and linear attention and state space models have become popular for improving long-context efficiency. However, efficiently parallelizing long-sequence training for recurrent and hybrid models remains challenging. We present StateFlow, a sequence pipeline parallelism system for models with linear recurrence. StateFlow partitions each sequence into chunks and schedules their execution while propagating boundary states and gradients across chunks, thereby reducing activation lifetimes and improving training throughput. StateFlow further uses profile-guided nonuniform chunking to balance recurrence and softmax attention computation in hybrid models, and overlaps state transitions that expose limited parallelism with surrounding computation. Applying StateFlow to models with up to 32B parameters and 256K context length, we achieve up to \(2.22\times\) throughput improvements and \(2.45\times\) memory reduction compared to conventional pipeline parallelism, enabling otherwise infeasible configurations.

↑