发表机构
Hunan University; China Unicom Software Research Institute; China Unicom Research Institute; Shanghai Jiao Tong University(湖南大学; 中国联通软件研究院; 中国联通研究院; 上海交通大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对LLM训练中OCS调度因双工端口对导致容量闲置的问题,提出LACE,首个独立分配收发车道的离线调度编译器,在不改变集合通信算法下实现非对称连接,实验显示最高2.04倍通信加速。
AI 中文摘要
光路交换机(OCS)能够重新配置物理连接,以匹配大语言模型(LLM)训练的可预测通信调度。尽管每条OCS光路在物理上是单工的,现有的需求感知型OCS调度器却以双工端口对的方式分配容量,强制两个方向具有相等的带宽,从而在非对称节点对流量下造成容量闲置。本文提出了LACE,这是首个离线OCS调度编译器,它独立地为LLM训练分配发送(TX)和接收(RX)车道。在不改变所选集合通信算法、操作顺序或排名放置的情况下,LACE重构有向的节点级需求,联合确定哪些连续操作共享一个配置以及每条方向由多少条单工电路服务,并在每节点车道库存和多OCS结构约束下将这些分配实现为物理车道绑定和光路。软件确认通过独立配置的返回路径携带反馈,而协调的链路配置和恢复在通信恢复前验证每个配置。在采用固定拓扑和匹配的每端口速率限制的独立三服务器测试台上,LACE的非对称连接在通信重放中实现了$1.80\times$的加速,在GPT-2训练中实现了$1.27\times$的加速,相对于对称拓扑基线。在更大规模下,对LLaMA-3.1 70B和405B调度(每服务器十六个400-Gb/s端口)的模拟显示,LACE相对于最新的双工OCS调度器实现了$1.21$--$2.04\times$的通信加速。
英文摘要
Optical circuit switch (OCS) can reconfigure physical connectivity to match the predictable communication schedules of large language model (LLM) training. Although each OCS light path is physically simplex, existing demand-aware OCS schedulers allocate capacity in duplex-port pairs, forcing equal bandwidth in both directions and stranding capacity under asymmetric node-pair traffic. This paper present LACE, the first offline OCS schedule compiler that independently allocates transmit (TX) and receive (RX) lanes for LLM training. Without changing the selected collective algorithms, operation order, or rank placement, LACE reconstructs directed node-level demand, jointly determines which consecutive operations share a configuration and how many simplex circuits serve each direction, and realizes these allocations as physical lane bindings and optical paths under per-node lane-inventory and multi-OCS fabric constraints. Software acknowledgments carry feedback over independently provisioned return paths, while coordinated link configuration and recovery verify each configuration before communication resumes. On a separate three-server testbed using fixed topologies and matched per-port rate limits, LACE's asymmetric connectivity achieves $1.80\times$ speedup for communication replay and $1.27\times$ for GPT-2 training over a symmetric-topology baseline. At larger scale, simulations of LLaMA-3.1 70B and 405B schedules with sixteen 400-Gb/s ports per server show that LACE achieves $1.21$--$2.04\times$ communication speedup over the latest duplex OCS scheduler.
Comments17 pages, 12 figures