发表机构
Korea University(高丽大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
SwiftQK是一种多GPU的RMSNorm核,通过仅交换标量归一化统计量并重叠归约与计算,大幅降低了张量并行下查询-键归一化的通信开销与延迟,提升了大型语言模型的推理效率。
AI 中文摘要
查询-键归一化(QK-Norm)可提升现代大型语言模型(LLM)的训练稳定性与质量,但在张量并行(TP)场景下,逐层QK-Norm会因归一化因子依赖完整隐藏向量而引入额外的跨GPU通信。本文提出SwiftQK,一种多GPU的RMSNorm核,仅交换标量归一化统计量,并将剩余点对点归约与独立元素级计算在死锁安全的持久核中重叠执行。对近期LLM的评估显示,相较于采用全向量All-Gather的标准TP QK-Norm,SwiftQK可将QK-Norm延迟降低81.4%至93.9%;在端到端推理中,相较于基于All-Gather的基线,SwiftQK平均降低端到端时间(TPOT)29.5%,相较于优化后的标量聚合实现则降低14.3%。
英文摘要
Query-Key Normalization (QK-Norm) improves the training stability and quality of modern Large Language Models (LLMs). However, under Tensor Parallelism (TP), layerwise QK-Norm introduces additional cross-GPU communication because the normalization factor depends on the full hidden vector. We present SwiftQK, a multi-GPU RMSNorm kernel that exchanges only scalar normalization statistics and overlaps the remaining Peer-to-Peer reduction with independent element-wise computation in a deadlock-safe persistent kernel. Evaluations on recent LLMs show that SwiftQK reduces QK-Norm latency by 81.4--93.9% relative to the standard TP QK-Norm using full-vector All-Gather. In end-to-end serving, SwiftQK reduces TPOT on average by 29.5% over the All-Gather-based baseline and by 14.3% over an optimized scalar-aggregation implementation.