AI 中文总结
针对数据中心网络微突发与高并发问题,提出概率状态比例(PSP)报文级负载均衡算法,其性能优于 JSQ、随机调度,与 Top-k 相当且硬件成本更低,实现性能、稳定性与开销的平衡。
AI 中文摘要
随着大语言模型训练与生成式人工智能服务的快速发展,数据中心网络面临严重的微突发流量与高并发问题。传统基于哈希的流级负载均衡无法感知链路状态,导致多路径 Clos 网络中出现哈希冲突、热点拥塞与尾部延迟。现有报文级方案受限于 stale 状态信息、高硬件复杂度及对异构链路适配性差的问题。为解决这些问题,本文提出概率状态比例(PSP)调度算法,这是一种报文级负载均衡算法。PSP 采用基于带宽的离散状态表示,用局部概率映射替代全局排序,降低硬件复杂度的同时抑制 stale 状态引发的拥塞聚集与振荡。在周期精确模拟器上的实验表明,PSP 在端口规模、带宽受限路径及固定流干扰场景下均具有鲁棒性,其在丢包率、99 百分位缓冲区占用率及可扩展性方面均优于最短队列优先(JSQ)调度与随机调度,且在硬件成本较低的情况下与 Top-k 调度性能相当。PSP 为人工智能数据中心提供了性能、稳定性与开销之间的有效平衡。
英文摘要
With the rapid growth of large language model training and generative artificial intelligence services, data center networks face severe micro-burst traffic and high concurrency. Traditional hash-based flow-level load balancing cannot sense link states, leading to hash collisions, hotspot congestion, and tail latency in multipath Clos networks. Existing packet-level schemes are constrained by stale state information, high hardware complexity, and poor adaptation to heterogeneous links. To address these issues, this paper proposes probabilistic state-proportional (PSP) dispatching, a packet-level load balancing algorithm. Using a Band-based discrete state representation, PSP replaces global sorting with local probability mapping, reducing hardware complexity while suppressing herding and oscillations caused by stale states. Experiments on a cycle-accurate simulator show that PSP is robust across port scales, bandwidth-limited paths, and fixed-flow interference. It outperforms join-the-shortest-queue (JSQ) scheduling and Random in loss rate, 99th-percentile buffer occupancy, and scalability, while remaining competitive with Top-k at lower hardware cost. PSP provides an effective balance among performance, stability, and overhead for artificial intelligence data centers.
Comments12 pages, 10 figures, 5 tables. Accepted by IEEE LCN 2026