arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.17644cs.DC

MLA序列并行中的训练-内存回归:为什么Megatron-Core禁止吸收,以及LAGA——一种通信高效的修复方法

A Training-Memory Regression in MLA Sequence Parallelism: Why Megatron-Core Forbids Absorption, and LAGA -- a Communication-Efficient Fix

Changzheng Ma

首次发表
浏览论文内容

中文总结 AI 辅助

研究MLA在Megatron-Core中训练时内存问题及限制原因,提出LAGA方法,该方法保留通信优势并解决内存问题,在特定硬件上减少通信、提升吞吐量,解决了无低通信量MLA训练路径的问题。

中文摘要 AI 辅助

多头潜在注意力(MLA)在Megatron-Core中有两种实现方式:一种用于训练的显式形式,另一种是吸收形式(通过仅收集压缩潜在值来大幅减少集合通信),但该吸收形式在训练中完全实现却被硬断言禁止(前向传播开头为“assert not (this http URL and self.cache_mla_latents)”),仅在推理解码中允许。库文档未说明原因。我们证明了该限制是有充分依据的,并量化了原因:移植到训练中时,吸收形式是一个内存陷阱,其每个令牌的中间值存在于n_h x d_kv维度,大于它们所取代的每个头的K/V,使激活内存增加20%-34%,在DeepSeek-V3规模下高达9.2GB(n_h = 128,seq = 1638, SP = 8,急切内核;在融合内核下差距扩大到19.2GB),足以改变设备适配情况。此测量在两个轴上得到验证(与seq和n_h呈线性关系)并在NVIDIA A100上交叉验证,解释了这一未记录的限制,使从业者没有低通信量的MLA训练路径。然后我们提供了一种方法。LAGA(潜在全收集注意力)保留了吸收形式的潜在收集通信,但拒绝吸收重新公式化,而是从收集的潜在值中本地重建每个头的K/V。在实际DeepSeek-V3维度的8x Ascend 910B上,LAGA将集合通信减少1.98倍,内存与显式形式相差在0.5%以内,在SP = 1时与显式形式位相同,在SP = 2 - 8时等效性在1e-3以内,并且在融合注意力内核下,单节点注意力块吞吐量提高1.04 - 1.06倍,跨节点提高1.07 - 1.24倍——在跨节点模式下所有序列长度中领先,而MLA正是为此模式部署的。

英文摘要

Multi-head Latent Attention (MLA) ships two implementations in Megatron-Core: an explicit form used for training and an absorbed form -- which slashes collective communication by gathering only the compressed latent -- that is fully implemented but hard-asserted out of training (the forward opens with "assert not (self.training and self.cache_mla_latents)"), allowed only in inference decode. The library documents no reason. We show the restriction is well-founded and quantify why: ported to training, the absorbed form is a memory trap -- its intermediates live in n_h x d_kv dimensions per token, larger than the per-head K/V they replace -- inflating activation memory by 20-34%, up to 9.2 GB at DeepSeek-V3 scale (n_h=128, seq=16384, SP=8, eager kernel; the gap widens to 19.2 GB under a fused kernel), enough to change device-fit. This measurement, validated on two axes (linear in seq and n_h) and cross-verified on NVIDIA A100, explains the otherwise-undocumented restriction and leaves practitioners with no low-communication MLA training path. We then provide one. LAGA (Latent All-Gather Attention) keeps the absorbed form's latent-gather communication but rejects the absorb reformulation, instead reconstructing per-head K/V locally from the gathered latent. On 8x Ascend 910B at real DeepSeek-V3 dimensions, LAGA cuts collective communication 1.98x, matches explicit memory within 0.5%, is bit-identical to explicit at SP=1 and equivalent to within 1e-3 at SP=2-8, and under a fused attention kernel improves attention-block throughput 1.04-1.06x single-node and 1.07-1.24x cross-node -- leading at all sequence lengths in the cross-node regime MLA is deployed for.

↑