发表机构
University of Chinese Academy of Sciences; Alibaba Group; National University of Singapore(中国科学院大学; 阿里巴巴集团; 新加坡国立大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究长上下文语言模型训练的负载平衡问题,提出Libra方法,利用大数定律,通过有界序列池、方差减少序列放置和平铺注意力池化等技术,有效减少注意力工作负载倾斜,提高训练吞吐量。
AI 中文摘要
长上下文语言模型训练存在负载平衡问题,序列打包无法解决。将样本打包到固定令牌序列中可平衡内存和线性成本算子,但注意力成本随序列长度平方和缩放。来自长尾语料库的等长打包序列可能承载不同注意力工作负载,导致数据并行掉队者和流水线气泡。现有方法要么在序列或微批次粒度上平衡,要么在全局工作池上分散注意力,通信域随数据并行度增长。我们提出天秤座,将大数定律作为负载平衡的缩放原则:注意力平衡池无需随数据并行度增长。天秤座将打包序列及其计算分区组分组到固定大小的序列池中。随着数据并行扩展,天秤座添加池而非扩大每个池,限制每次注意力交换。方差减少序列放置通过将具有互补注意力工作负载的序列共定位来减少池间剩余倾斜,对有限的长尾工作负载有效。在每个池内,平铺注意力池化在GPU之间调度序列头切片,流水线运行时重叠切片交换和注意力。天秤座公开了一个即插即用的上下文并行注意力算子和一个可插拔的数据采样器,无需对模型层、优化器或流水线调度进行更改。在使用256K和1M令牌工作负载的Qwen3-Turbo训练中,天秤座比尤利西斯提高了高达2.54倍的端到端吞吐量,在微基准测试中最差步骤掉队者注意力加速高达3.14倍。天秤座已在生产中运行了数十万个GPU小时,处理32K到1M令牌的任务,同时保留训练语义。
英文摘要
Long-context LLM training suffers from a load-balancing problem that sequence packing does not solve. Packing samples into fixed-token sequences balances memory and linear-cost operators, but the dominant attention cost scales with the sum of squared sequence lengths. Thus, equally sized packed sequences drawn from a long-tailed corpus can carry substantially different attention workloads, creating data-parallel stragglers and pipeline bubbles. Existing approaches either balance at the granularity of sequences or microbatches, where an outlier can dominate an assignment, or disaggregate attention over a global worker pool whose communication domain grows with the data-parallel (DP) degree. We present Libra, which operationalizes the law of large numbers (LLN) as a scaling principle for load balancing: the attention-balancing pool need not grow with the DP degree. Libra groups packed sequences and their CP groups into fixed-size sequence pools. As DP scales out, Libra adds pools rather than enlarging each one, bounding every attention exchange. Variance-Reduced Sequence Placement makes this effective for finite, long-tailed workloads by co-locating sequences with complementary attention workloads to reduce residual inter-pool skew. Within each pool, Tiled Attention Pooling dispatches sequence-head SH-Tiles across GPUs, while a pipelined runtime overlaps tile exchange with attention. Libra exposes a drop-in context-parallel attention operator and a pluggable data sampler, requiring no changes to model layers, optimizers, or pipeline schedules. On three production Qwen3 models (8B, 30B, 235B) and 256K- and 1M-token production workloads, Libra improves end-to-end training throughput over the strongest evaluated baseline (WLB-LLM) by 44% on average and up to 68% at 256 GPUs. Libra has run for hundreds of thousands of GPU-hours in production on jobs spanning 32K to 1M tokens.
Comments15 pages, 15 figures