arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

Cnuas:一个软件定义的AI/HPC机架级仿真平台与超大规模数据中心设施孪生

Cnuas: A Software-Defined AI/HPC Rack-scale Emulation Platform and Hyperscale Data Center Facility Twin

Weqaar Janjua, Eoin OConnell, Mihai Penica

arXiv 2609.15889首次发表:更新:

发表机构

University of Limerick; Packet Five Networks Ltd.(利默里克大学; Packet Five Networks有限公司)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出Cnuas,一个遵循OCP Open Rack v3规范的开源机架级仿真平台,通过功能仿真支持AI/HPC软件研发,并提供RDMA、交换机及OpenBMC管理接口,以促进社区协作。

AI 中文摘要

现代AI和高性能计算(HPC)系统在机架级集成了加速器、高速网络和管理控制器。为此类基础设施开发软件通常需要访问稀缺且昂贵的硬件,而软件抽象可能掩盖工作负载对跨服务器和加速器资源依赖的方式。本文介绍了Cnuas,一个开源的、实验性的机架级仿真平台,其基线架构遵循开放计算项目(OCP)Open Rack v3规范。通过功能仿真,它支持学术和工业研发中的实验、学习及软件开发,而非匹配物理硬件的吞吐量或延迟。其基于Web的用户界面可视化机架、设备及其互连,帮助开发者建立支撑其工作负载的基础设施的系统级心智模型。其核心是,CnuasNIC和CnuasSwitch实现了一个客户可见的远程直接内存访问(RDMA)适配器和一个主机驻留的混合软件交换机,同时支持RoCEv2和原生InfiniBand。该平台还提供了一个专用的AI/ML加速器(GPU)对等结构,以及基于OpenBMC的机架管理,带有通过RS-485的可执行电源和电池备份固件。这些组件支持在商用主机上研究设备、驱动和固件接口。加速器软件栈仍处于早期研究原型阶段,使用OpenUSD进行设施建模是一个探索性扩展。本文介绍了架构、接口和有界原型结果,作为围绕核心平台及其扩展进行社区协作的基础。

英文摘要

Modern AI and HPC systems integrate accelerators, high-speed networks, and management controllers at rack scale. Developing software for this infrastructure typically requires access to scarce, costly hardware, while software abstractions can obscure how workloads depend on resources across servers and accelerators. This paper presents Cnuas, an open-source, experimental rack-scale emulation platform whose baseline architecture follows the Open Compute Project (OCP) Open Rack v3 specifications. Through functional emulation, it supports experimentation, learning and software development within academic and industrial research and development, rather than matching the throughput or latency of physical hardware. Its web-based user interface visualizes racks, devices and their interconnections to help developers build a system-level mental model of the infrastructure supporting their workloads. At its core, CnuasNIC and CnuasSwitch implement a guest-visible remote direct memory access (RDMA) adapter and a host-resident hybrid software switch supporting both RoCEv2 and native InfiniBand. The platform also provides a dedicated AI/ML accelerator (GPU) peer fabric and OpenBMC-based rack management with executable power supply and battery backup firmware over RS-485. These components support the study of device, driver, and firmware interfaces on commodity hosts. The accelerator software stack remains an early research prototype, and facility modeling with OpenUSD is an exploratory extension. The paper presents the architecture, interfaces, and bounded prototype results as a basis for community collaboration across the core platform and its extensions.

Comments13 pages, 5 figures, 4 tables

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑