arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

MoSE:用于混合LLM训练与推理的模式切换扩展器

MoSE: Mode-Switching Expander for Mixed LLM Training and Inference

Fan Yang, Ying Zhou, Binglei Wang, Zhenjie Zhou, Jialong Li

arXiv 2609.39138首次发表:更新:

发表机构

Southern University of Science and Technology; Faculty of Computer Science and Artificial Intelligence, Shenzhen University of Advanced Technology; School of Electronic and Information Engineering, Beijing Jiaotong University(南方科技大学; 深圳先进技术大学计算机科学与人工智能学院; 北京交通大学电子与信息工程学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

MoSE是一种可重构扩展器,通过将拓扑设计视为固定度数的边分配问题,在推理密集型模式下优化P-D连接,在训练密集型模式下恢复均匀随机正则扩展器,显著降低KV通信和训练通信成本。

AI 中文摘要

AI集群日益在同一网络上运行大型语言模型(LLM)的推理和训练任务。预填充-解码(P-D)分离架构在预填充组和解码组之间产生键值(KV)缓存传输,而训练集合通信和全对全流量则受益于接近均匀的全局连接。因此,静态稀疏拓扑可能无法很好地匹配这两种流量模式之一。我们提出了模式切换扩展器(MoSE),一种可重构的扩展器,将拓扑设计视为固定度数的边分配问题。MoSE将相同的稀疏边预算重新分配,在推理密集型模式下用于直接的P-D连接,并在训练密集型模式下恢复为均匀随机正则扩展器。我们使用1024组流级拓扑模型、最短路径路由和两种混合工作负载对MoSE进行了评估。在20个随机种子中,相对于静态训练拓扑,MoSE在推理密集型模式下将平均和95分位(P95)负载感知的KV通信成本分别降低了90.8%和91.9%。在训练密集型模式下,相对于过时的静态推理拓扑,MoSE将平均和P95训练通信成本分别降低了22.7%和27.6%。这些结果表明,粗粒度的拓扑切换可以在不增加端口或更改路由的情况下支持两种工作负载模式。

英文摘要

AI clusters increasingly run large language model (LLM) inference and training on the same fabric. Prefill-decode (P-D) disaggregation creates key-value (KV) cache transfers between prefill and decode groups, whereas training collectives and all-to-all traffic benefit from near-uniform global connectivity. A static sparse topology can therefore be poorly matched to one of the two traffic patterns. We present Mode-Switching Expander (MoSE), a reconfigurable expander that treats topology design as a fixed-degree edge-allocation problem. MoSE reallocates the same sparse edge budget toward direct P-D connectivity in inference-heavy modes and restores a uniform random regular expander in training-heavy modes. We evaluate MoSE using a 1024-group flow-level topology model, shortest-path routing, and two mixed workloads. Across 20 seeds, MoSE reduces average and 95th-percentile (P95) load-aware KV communication cost by 90.8\% and 91.9\% relative to Static-Training in the inference-heavy mode. In the training-heavy mode, it reduces average and P95 training communication cost by 22.7\% and 27.6\% relative to stale Static-Inference. These results show that coarse-grained topology switching can support both workload modes without additional ports or routing changes.

CommentsAccepted by ACP 2026, Top-Score Papers

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑