arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

分布式AI推理何时需要更多广域网带宽?光学、分组与软件杠杆的协同设计评估

When Does Distributed AI Inference Need More Wide-Area Bandwidth? A Co-Design Evaluation of Optical, Packet, and Software Levers

Prasanna C

arXiv 2608.14967首次发表:更新:

AI 中文总结

本文对比了分布式AI推理中,跨站点迁移推理状态与KV重计算等方案的优劣,推导了70B多头注意力模型的带宽交叉点,分析了敏感性因素及经济性,提出了测试计划。

AI 中文摘要

每一代硬件中,单位GPU算力对应的广域网带宽不断下降:以计算强度比(CIR,每浮点运算对应的字节数)衡量,封装内内存与传统广域网之间的差距达4至5个数量级,且每年以约12%至19%的幅度扩大。此前的立场文件(包括我们自己的)认为,这使得跨站点AI推理需要弹性光学广域网容量。审稿人正确地指出,这类论点仅表明更多带宽有帮助,并未证明其优于替代方案:KV重计算、缓存压缩、感知局部性的路由与调度,或过度配置的分组骨干网。本文完成了这一对比。我们推导了一个工作负载模型,用于预测何时跨站点迁移推理状态比重计算更具优势——对于700亿参数的多头注意力模型,每条流的上下文无关交叉点为74至111 Gbps,在分组查询注意力(分组查询注意力)下,当KV头数量为1/8时,该值降至9至14 Gbps;我们还量化了五个敏感性维度:上下文长度、注意力架构、排队、智能体复合以及丢包/抖动导致的带宽崩溃。在经济性方面:按GPU挂牌价格计算,重计算更便宜;当GPU稀缺性与KV复用将有效GPU成本提高约5至20倍时,传输更具优势,而现代注意力机制使传输的盈亏平衡点向有利于传输的方向移动了一个数量级。我们正确定位网络杠杆而非采取对抗态度:分组网络在毫秒级时间尺度分配已点亮的容量;光学可替代性改变了点亮容量的多少,时间尺度为分钟级,在经济上而非功能上替代过度配置。最后,我们在三站点生产光纤测试台上制定了十项指标的测量计划,将其表述为开放生态系统工作:没有任何一家公司能够或应当单独收集这些证据。所有结论均受其适用范围限制;部分发现削弱了我们自身论点的朴素版本,我们明确指出这些限制。

英文摘要

Wide-area bandwidth per unit of GPU compute falls every hardware generation: in compute-intensity-ratio terms (CIR, bytes per FLOP), the gap between on-package memory and the conventional WAN is four to five orders of magnitude, widening at roughly 12-19% per year. Position papers - including our own - argued this makes elastic optical wide-area capacity necessary for cross-site AI inference. Reviewers correctly objected that such arguments show more bandwidth helps, not that it beats the alternatives: KV recomputation, cache compression, locality-aware routing, scheduling, or an overprovisioned packet backbone. This paper does the comparison. We derive a workload model predicting when moving inference state across sites beats recomputing it - a context-independent crossover at 74-111 Gbps per stream for a 70B multi-head-attention model, falling to 9-14 Gbps under grouped-query attention at 1/8 KV heads - and quantify five sensitivity axes: context length, attention architecture, queueing, agentic compounding, and loss/jitter-induced bandwidth collapse. On economics: at list GPU prices recomputation is cheaper; transfer wins when GPU scarcity and KV reuse multiply effective GPU cost by roughly 5-20x, and modern attention moves the breakeven an order of magnitude in transfer's favour. We position the network levers correctly rather than adversarially: packet networks allocate lit capacity at millisecond timescales; optical fungibility changes how much capacity is lit, at minute timescales, substituting for overprovisioning economically rather than functionally. Finally we specify a ten-metric measurement plan on a three-site production-fibre testbed, framed as an open ecosystem exercise: no single company can - or should - assemble this evidence alone. Every claim is bounded by the regime in which it holds; several findings weaken the naive version of our own thesis, and we state them.

Comments10 pages, 4 figures, 3 tables

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑