arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.21893cs.NI

路径越少,性能越好:通过布雷斯悖论理解 ZCube 拓扑结构

Fewer Paths, Better Performance: Understanding the ZCube Topology through Braess's Paradox

Li Chen

首次发表
浏览论文内容

中文总结 AI 辅助

研究数据中心网络中 ZCube 拓扑结构,通过与布雷斯悖论的联系,揭示其在模型训练和推理中性能更好的原因,证明其不受悖论影响且能平衡负载,还量化了代价,指出拓扑与流量匹配比路径多样性更重要。

中文摘要 AI 辅助

数据中心网络遵循多路径原则,而 ZCube 拓扑结构违反了这一原则,它移除了脊柱层,消除了路径多样性并削减了三分之一的交换硬件,但在大型模型训练和推理方面却具有更好的性能。我们通过与 1968 年首次观察到的布雷斯悖论的结构联系来解释这一异常现象。首先,我们表明在结构化语言模型流量下的多路径结构在布雷斯阴影下运行,静态 ECMP 哈希比贪婪路由更脆弱。其次,我们证明 ZCube 不受布雷斯悖论影响,其正交对偶分区可证明能平衡任意流量矩阵的负载。第三,我们量化了这种免疫的代价。来自服务 GLM - 5.1 编码推理的集群的生产测量报告显示网络成本降低 33%,GPU 吞吐量提高 15%,首次令牌的 P99 时间降低 40.6%。我们的分析表明,对于由模型结构驱动的工作负载,使拓扑结构与流量匹配比路径多样性更重要。

英文摘要

Datacenter networks follow a multipath doctrine: provision many paths between endpoints, hash flows across them, and let redundancy absorb both failures and load imbalance. The ZCube topology violates this doctrine. It removes the Spine layer, eliminates path multiplicity, and cuts one third of switching hardware, yet delivers better performance for both large model training and inference. We explain this anomaly through a structural connection to Braess's paradox, first observed in 1968: both phenomena trace to congestion-oblivious routing over competing paths. Braess showed that adding paths under this condition can hurt; ZCube shows that removing paths under the same condition can help. First, we show that multipath fabrics under structured LLM traffic operate in Braess's shadow: static ECMP hashing is strictly more fragile than greedy routing. Greedy routing reaches an equilibrium within 4/3 of optimal for affine latencies; static hashing admits unbounded imbalance in the worst case. Second, we prove that ZCube is immune to Braess's paradox and that its orthogonal dual partition provably balances load for arbitrary traffic matrices; AllReduce in training and KV cache transfers in disaggregated inference fall out as two corollaries. Third, we quantify the price of this immunity: ZCube trades microsecond hash recovery for millisecond control plane recovery, a trade that upper-layer resilience in LLM serving makes favorable. Production measurements from a cluster serving GLM-5.1 coding inference report 33% lower network cost, 15% higher GPU throughput, and 40.6% lower P99 time to first token. Our analysis suggests that for workloads driven by model structure, matching topology to traffic matters more than path multiplicity.

补充信息

↑