KV-Pipe:LLM中KV共享与流水线并行效率的关系
KV-Pipe: On the Relation Between KV Sharing and Pipeline Parallel Efficiency in LLMs
浏览论文内容
中文总结 AI 辅助
KV-Pipe是一种感知阶段的KV共享机制,通过尾优先调整跨层KV共享平衡流水线工作负载,提升LLM训练效率,同时也优化长上下文推理性能。
中文摘要 AI 辅助
流水线并行(PP)被广泛用于扩展大语言模型(LLM)的训练,但其效率常受阶段不平衡和流水线气泡的限制。同时,跨层KV共享主要被作为推理期间降低KV缓存成本的机制研究,未探究KV复用如何重塑流水线工作负载。我们提出了KV-Pipe,一种感知阶段的KV共享机制,可将KV复用转化为流水线平衡的控制旋钮。KV-Pipe从尾阶段开始,以尾优先顺序将选定的注意力层转换为跨层KV共享,并迭代重定向当前瓶颈以推动FLOPs不平衡比率(FIR)趋近于1。该过程离线执行,仅需流水线划分和每层FLOPs估计,引入的运行时开销可忽略,且无需在线调优。在多个流水线并行配置中,KV-Pipe持续提升利用率和吞吐量,实现训练MFU最高提升9.2%,迭代时间最高减少9.8%,在阶段不平衡被放大的更高流水线并行度下增益更大。此外,相同的KV共享机制在推理侧也有好处,可减少KV缓存增长和冗余KV投影工作,为长上下文工作负载带来更高的解码吞吐量。这些结果表明KV布局是可协同提升流水线并行训练效率和长上下文推理的系统-架构自由度。
英文摘要
Pipeline parallelism (PP) is widely used to scale large language model (LLM) training, but its efficiency is often limited by stage imbalance and pipeline bubbles. Meanwhile, cross-layer KV sharing has primarily been studied as a mechanism for reducing KV-cache costs during inference, without examining how KV reuse reshapes pipeline workloads. We present \textbf{KV-Pipe}, a stage-aware KV-sharing mechanism that turns KV reuse into a pipeline-balancing control knob. KV-Pipe starts from the tail stage, converts selected attention layers to cross-layer KV sharing in a tail-first order, and iteratively retargets the current bottleneck to drive the FLOPs Imbalance Ratio (FIR) toward $1$. The procedure is performed offline and requires only a pipeline partition and per-layer FLOPs estimates, introducing negligible runtime overhead and requiring no online tuning. Across multiple pipeline-parallel configurations, KV-Pipe consistently improves utilization and throughput, achieving up to \textbf{9.2\%} higher training MFU and up to a \textbf{9.8\%} reduction in iteration time, with larger gains at higher pipeline-parallel degrees where stage imbalance is amplified. Furthermore, the same KV-sharing mechanism provides an inference-side benefit by reducing KV-cache growth and redundant KV projection work, resulting in higher decoding throughput for long-context workloads. These results identify KV layout as a system--architecture degree of freedom for jointly improving pipeline-parallel training efficiency and long-context inference.