arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

面向延迟张量并行的亲和感知分片

Affinity-Aware Sharding for Delayed Tensor Parallelism

Eloi de Reynal

arXiv 2609.13846首次发表:更新:

AI 中文总结

针对延迟张量并行,通过置换模型最大化KV头与FFN神经元共置亲和性,加速蒸馏重训练,实验显示所需步数减半至三分之二。

AI 中文摘要

延迟张量并行(DTP)移除了张量并行Transformer推理中的阻塞式全归约。每个设备立即将其部分输出添加到其残差流中(并广播),但仅在δ个模块之后才收集(接收)其他设备的部分输出。因此,从TP到DTP的转变相当于一次真正的架构变更,密集Transformer模型在适配后需要重新训练或蒸馏。我们证明DTP打破了FFN内部神经元和注意力模块内部KV头的置换对称性,而这种对称性破缺使得分片本身成为一个建模决策。我们表明,通过在分片之前对密集模型进行置换,最大化同一设备上共置的KV头与FFN神经元之间的亲和性,可以加速蒸馏或重新训练过程。亲和性通过一阶近似来衡量丢失某个头的贡献对每个神经元输出的损害,共置亲和性通过坐标上升优化器最大化,该优化器交替进行神经元的精确平衡分配和对KV头划分的穷举搜索。整个过程在单个GPU上对Qwen3-0.6B和Danube3-500M耗时不到两分钟。在这些模型上,当δ=1时,在我们测试的整个10k步范围内,亲和性优化的布局达到任何蒸馏目标所需的步数约为朴素连续布局的一半至三分之二,并且每个优化的随机种子都优于每个连续种子,且优于十六个随机布局中的十五个。我们还表明,初始化时的共置亲和性分数可以预测训练后与基础模型的KL散度,涵盖从反优化到优化的十七种布局(Pearson相关系数分别为-0.81和-0.89)。

英文摘要

Delayed Tensor Parallelism (DTP) removes the blocking all-reduce of tensor-parallel Transformer inference. Every device adds its own partial output to its residual stream (and broadcasts it) immediately, but only gathers (receives) the other devices' partials $δ$ modules later. A TP to DTP change therefore amounts to a real architecture change, and dense Transformer models need to be retrained or distilled after adaptation. We show that DTP breaks the permutation symmetry of neurons inside FFNs and of KV heads inside attention modules, and that this symmetry breakage makes the sharding itself a modelling decision. We show that maximising the affinity between the KV heads and the FFN neurons co-located on a device, by permuting the dense model before sharding, speeds up the distillation or retraining process. The affinity is measured with a first-order approximation of the damage that losing a head's contribution does to each neuron's output, and the co-located affinity is maximised with a coordinate-ascent optimiser that alternates an exact balanced assignment of neurons with an exhaustive search over the KV head partitions. The whole procedure takes under two minutes on one GPU for Qwen3-0.6B and Danube3-500M. On these models, at $δ=1$, the affinity-optimised layouts reach any distillation target in about half to two thirds of the steps needed by the naive contiguous layouts, over the whole 10k-step range we tested, and every optimised seed beats every contiguous seed and all but one of the sixteen random layouts. We also show that the co-located affinity score at initialisation predicts the KL to the base model after training, across seventeen layouts ranging from anti-optimised to optimised (Pearson $-0.81$ and $-0.89$).

Comments16 pages, 11 figures, 3 tables

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑