WaferTrans:为晶圆级GPU实现无IOMMU的分布式虚拟地址翻译
WaferTrans: Enabling IOMMU-free Distributed Virtual Address Translation for Wafer-scale GPUs
浏览论文内容
中文总结 AI 辅助
WaferTrans提出无IOMMU的分布式虚拟地址翻译方案,通过PTE存在性一致性和目录机制,在晶圆内解析远程翻译,平均性能提升2.5倍。
中文摘要 AI 辅助
晶圆级GPU(WSG)提供了充足的晶圆上带宽,使得近乎无损的统一内存成为可能。然而,现有设计仍依赖CPU-IOMMU来翻译远程虚拟地址访问。这种集中式机制在扩展到数十个GPU裸片时表现不佳:翻译请求必须穿越昂贵的晶圆外层级结构,并争用有限的CPU侧资源,使得地址翻译成为关键瓶颈。我们提出了WaferTrans,一种面向WSG的无IOMMU分布式虚拟地址翻译设计。WaferTrans引入了PTE存在性一致性(PTE-PC),一种轻量级一致性模型,用于跟踪PTE的插入和移除,并为每个GPU配备一个PTE存在性目录(PPD),以定位持有所需PTE的GPU。它进一步采用分布式PTE-PC映射和协作查询机制,在保持完整查找覆盖的同时,将PTE-PC维护本地化。这些机制共同使GPU阵列能够在晶圆内解析远程翻译,消除了其对CPU-IOMMU的依赖。与最先进的Trans-FW设计相比,WaferTrans平均性能提升了2.5倍。
英文摘要
Wafer-scale GPUs (WSGs) provide sufficient on-wafer bandwidth to make near-lossless Unified Memory feasible. However, existing designs still rely on a CPU-IOMMU to translate remote virtual-address accesses. This centralized mechanism scales poorly to tens of GPU dies: translation requests must traverse costly off-wafer hierarchies and contend for limited CPU-side resources, making address translation a critical bottleneck. We propose WaferTrans, an IOMMU-free distributed virtual-address translation design for WSGs. WaferTrans introduces PTE Presence Consistency (PTE-PC), a lightweight consistency model that tracks PTE insertions and removals, and equips each GPU with a PTE Presence Directory (PPD) that locates the GPU holding a requested PTE. It further employs a distributed PTE-PC mapping and a cooperative query mechanism to localize PTE-PC maintenance while preserving complete lookup coverage. Together, these mechanisms enable the GPU array to resolve remote translations within the wafer, eliminating its dependence on the CPU-IOMMU. Compared with the SOTA Trans-FW design, WaferTrans improves performance by 2.5x on average.
发表机构
- School of Integrated Circuits and Systems, Tsinghua University(清华大学集成电路学院)
机构由 AI 辅助整理,请以论文原文为准。