arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.15943cs.DC

KV-Pipe:LLM中KV共享与流水线并行效率的关系

KV-Pipe: On the Relation Between KV Sharing and Pipeline Parallel Efficiency in LLMs

Maryam Dialameh, Hossein Rajabzadeh, Harish Krishnamoorthy Murali, Walid Ahmed, Weiwei Zhang, Hyock Ju Kwon

首次发表
浏览论文内容

中文总结 AI 辅助

KV-Pipe是一种感知阶段的KV共享机制,通过尾优先调整跨层KV共享平衡流水线工作负载,提升LLM训练效率,同时也优化长上下文推理性能。

中文摘要 AI 辅助

流水线并行(PP)被广泛用于扩展大语言模型(LLM)的训练,但其效率常受阶段不平衡和流水线气泡的限制。同时,跨层KV共享主要被作为推理期间降低KV缓存成本的机制研究,未探究KV复用如何重塑流水线工作负载。我们提出了KV-Pipe,一种感知阶段的KV共享机制,可将KV复用转化为流水线平衡的控制旋钮。KV-Pipe从尾阶段开始,以尾优先顺序将选定的注意力层转换为跨层KV共享,并迭代重定向当前瓶颈以推动FLOPs不平衡比率(FIR)趋近于1。该过程离线执行,仅需流水线划分和每层FLOPs估计,引入的运行时开销可忽略,且无需在线调优。在多个流水线并行配置中,KV-Pipe持续提升利用率和吞吐量,实现训练MFU最高提升9.2%,迭代时间最高减少9.8%,在阶段不平衡被放大的更高流水线并行度下增益更大。此外,相同的KV共享机制在推理侧也有好处,可减少KV缓存增长和冗余KV投影工作,为长上下文工作负载带来更高的解码吞吐量。这些结果表明KV布局是可协同提升流水线并行训练效率和长上下文推理的系统-架构自由度。

英文摘要

Pipeline parallelism (PP) is widely used to scale large language model (LLM) training, but its efficiency is often limited by stage imbalance and pipeline bubbles. Meanwhile, cross-layer KV sharing has primarily been studied as a mechanism for reducing KV-cache costs during inference, without examining how KV reuse reshapes pipeline workloads. We present \textbf{KV-Pipe}, a stage-aware KV-sharing mechanism that turns KV reuse into a pipeline-balancing control knob. KV-Pipe starts from the tail stage, converts selected attention layers to cross-layer KV sharing in a tail-first order, and iteratively retargets the current bottleneck to drive the FLOPs Imbalance Ratio (FIR) toward $1$. The procedure is performed offline and requires only a pipeline partition and per-layer FLOPs estimates, introducing negligible runtime overhead and requiring no online tuning. Across multiple pipeline-parallel configurations, KV-Pipe consistently improves utilization and throughput, achieving up to \textbf{9.2\%} higher training MFU and up to a \textbf{9.8\%} reduction in iteration time, with larger gains at higher pipeline-parallel degrees where stage imbalance is amplified. Furthermore, the same KV-sharing mechanism provides an inference-side benefit by reducing KV-cache growth and redundant KV projection work, resulting in higher decoding throughput for long-context workloads. These results identify KV layout as a system--architecture degree of freedom for jointly improving pipeline-parallel training efficiency and long-context inference.

↑