发表机构
Huawei Technologies Co., Ltd.(华为技术有限公司)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出嵌套BSP和嵌套并行冯·诺依曼架构,通过递归分层与统一总线实现百万级对等并行,使众多处理器协同为单一计算机。
AI 中文摘要
大规模AI计算不再是“单个更强处理器”的较量,而是关于一支受统一指挥的处理器大军如何仍然是一台计算机。本文提出两个相互关联的扩展。首先,将BSP扩展为嵌套BSP。图灵机将计算描述为单条磁带上的顺序操作。一百万个处理器不需要更长的磁带,而是需要一层嵌套一层的作战计划:在每一层,执行并行工作、屏障同步、交换与聚合,然后进入下一阶段。层内的每一次“并行推进”都重复这四个步骤。嵌套BSP通过递归嵌套经典BSP,成为一种面向百万级并行性的计算范式,其统一规则是:每一层的每个节点都是对等的。其次,将冯·诺依曼架构扩展为嵌套并行冯·诺依曼架构,统一总线作为其互连。冯·诺依曼教会我们如何构建一台存储程序计算机。八十年的错误外推是,将多台计算机连成网络就能得到一台更大的计算机。另一个更根深蒂固的习惯是:几乎每个设计都假设有一个发号施令的主机和服从的从机——主机支配设备,CPU支配加速器,中心支配边缘。嵌套并行架构扩展了这一思想而非抛弃它。两个嵌套娃娃必须契合:软件中的嵌套BSP,以及从封装到自治区域的嵌套并行冯·诺依曼架构,由一条端到端的内存语义总线连接,实现完全对等:物理上分散,逻辑上紧密。它与华为的τ缩放定律相辅相成:τ决定每一层如何折叠时间,而架构决定嵌套并行计算机如何逐层、逐对等地站立。总之,本文将BSP扩展为嵌套BSP,并将冯·诺依曼架构扩展为嵌套并行冯·诺依曼架构。τ折叠时间,对等并行性逐层嵌套——众多处理器,仍然是一台计算机。
英文摘要
The contest in large-scale AI computing is no longer whether one processor can be made stronger, but whether a million of them can still behave as a single computer. This paper argues that two extensions are required. First, classic BSP, nested recursively, becomes a plan that holds at every scale: each superstep consists of parallel work, then exchange and aggregation, then a barrier, and at every layer the participating units are peers. Second, the von Neumann single-machine architecture, extended past its master--slave habit, becomes the Nested Parallel von Neumann Architecture: a nest of peer-equal layers, from chip package to autonomous zone, joined end to end by one memory-semantic interconnect, the Unified Bus. The software nest and the hardware nest correspond layer by layer, and the $τ$ Scaling law folds time at every layer. We examine which workloads the nested structure serves, AI training above all but much of classic HPC as well, and describe the hardware decisions that make the nesting physical. Many processors, still one computer.
Comments14 pages, 4 figures