AI 中文总结
Shiftfly 用广义 Kautz 有向图替换 Boardfly 互连的全局光层级,在大规模(40 万芯片)下缩短最坏情况跳数,成本相同且频谱扩展提升约 2.7 倍,仅单 Pod 规模性能略差。
AI 中文摘要
谷歌的 TPU 互连历经九代均采用 k 元 n 立方体架构,其直径随 Θ(N^(1/n)) 增长,直到 TPU 8i 被 Boardfly 取代:Boardfly 是一个三层层级结构,其中每一层都是完全图,在 1152 个芯片的情况下,其 Pod 直径为 7。Boardfly 适用于单个推理 Pod,但无法扩展,因为完全全局层级需要 G-1 个光端口才能连接 G 个组。一个拥有 40 万个芯片的机器每个组需要 12499 个端口,而可用端口仅为 40 个,因此必须堆叠另一个层级,且每个层级会增加 4 次芯片跳数。我们提出 Shiftfly,它保留了 Boardfly 的基本构建块和组不变,仅将全局层级替换为广义 Kautz 有向图,并以固定排列形式部署在 fabric 已有的光电路交换机上。在每组相同的 40 个端口下,Shiftfly 是扁平结构,具有保证的直径 ⌈log_d G⌉,通过移位寄存器无需路由表即可路由,且原生处理内容。评估采用双向方式:Shiftfly 在单 Pod 规模上处于劣势,此时 Boardfly 的芯片级直径为 7,而 Shiftfly 为 11;但在更大规模上占优,在 40 万个芯片时将最坏情况距离从 23 跳缩短至 19 跳,且在成本相同的情况下频谱扩展性能提升约 2.7 倍。驱动该设计的冗余性论点在评估中不成立:感知局部性的放置方式提供了共享内容可实现的大部分节省,仅留下 2.9% 的剩余节省来自移位代数,我们还报告了掩盖这一情况的指标反转。我们还测量了可操作性:替换故障组在两种设计中都需要相同的 40 个光电路;推导全局布线所需的控制平面状态为 28 位,而 Boardfly 为 55 万位;切片分配是唯一的退化点:移位图的任意诱导子集是不连通的,因此切片必须实例化而不是切割。
英文摘要
Google's TPU interconnect spent nine generations as a $k$-ary $n$-cube, whose diameter grows as $Θ(N^{1/n})$, before TPU 8i replaced it with Boardfly: a three-tier hierarchy in which every tier is a complete graph, giving pod diameter 7 over 1,152 chips. Boardfly suits a single inference pod but does not extend, because a complete global tier needs $G-1$ optical ports to reach $G$ groups. A 400,000-chip machine would need 12,499 per group against the 40 available, so another hierarchy level must be stacked, and each level costs four chip hops. We propose Shiftfly, which keeps Boardfly's building block and group verbatim and replaces only the global tier with a generalized Kautz digraph, installed as a fixed permutation on the optical circuit switch the fabric already owns. At an identical 40 ports per group, Shiftfly is flat, has guaranteed diameter $\lceil \log_d G \rceil$, routes without tables by a shift register, and addresses content natively. The evaluation is deliberately two-sided. Shiftfly loses at one-pod scale, where Boardfly achieves chip-level diameter 7 against Shiftfly's 11, and wins beyond it, cutting worst-case distance from 23 to 19 hops at 400,000 chips with roughly 2.7x better spectral expansion at equal cost. The redundancy argument that motivated the design does not survive its own evaluation: locality-aware placement supplies most of the achievable saving on shared content, leaving the shift algebra a 2.9% residual, and we report the metric inversion that conceals this. We also measure operability. Replacing a failed group costs the same 40 optical circuits in both designs; deriving the global wiring costs 28 bits of control-plane state against 550,000. Slice allocation is the one regression: an arbitrary induced subset of a shift graph is disconnected, so slices must be instantiated rather than carved.
Comments8 pages, 5 figures, 5 tables. Simulator, proofs and reproduction artefacts: https://github.com/EylonKrause/shiftfly