发表机构
Clemson University(克莱姆森大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对现有序列并行方法通信效率低的问题,提出BASP技术,通过按微批次大小划分GPU减少通信开销,在Llama和Qwen模型上提升训练时间1.17-1.31倍且不损失精度与内存使用。
AI 中文摘要
大语言模型(LLM)的长上下文推理日益重要,但长序列训练因内存与通信需求巨大仍具挑战性。序列并行是解决长序列LLM训练瓶颈的关键技术,但现有序列并行方法与批次无关,对所有批次大小采用统一的序列划分,导致通信效率低下。本文提出批次感知序列并行(Batch-Aware Sequence Parallelism,BASP),一种利用批次结构降低通信开销的序列并行方法。BASP根据微批次大小将GPU划分为不相交的序列并行组,该设计减小了全对全通信组的规模,从而实现通信本地化并提升训练效率。在NVIDIA A100集群上的实验结果显示,与标准序列并行基线相比,BASP在Llama和Qwen模型上可将端到端训练时间提升至1.17-1.31倍,同时保持相同的模型精度和内存使用量。
英文摘要
Long-context reasoning for large language models (LLMs) is becoming increasingly important, but training over long sequences remains challenging due to massive memory and communication requirements. Sequence parallelism has emerged as an essential technique for addressing bottlenecks in long sequence LLM training. However, we observe that existing sequence parallelism methods are batch-agnostic and apply uniform sequence partitioning across all batch sizes, resulting in inefficient communication. In this paper, we introduce Batch- Aware Sequence Parallelism (BASP), a sequence parallelism approach that leverages batch structure to reduce communication overhead. BASP exploits batch structure by partitioning GPUs into disjoint sequence-parallel groups according to the micro- batch size. This design reduces the all-to-all communication group size, thereby localizing communication and improving training efficiency. Experimental results on an NVIDIA A100 cluster show that BASP improves end-to-end training time by up to 1.17 - 1.31x in Llama and Qwen models compared to standard sequence parallel baselines, while preserving identical model accuracy and memory usage.
CommentsAccepted by IEEE HPEC 2026