AI 中文总结
该研究针对AI研究中存储系统作用常被忽视的问题,引入PRISM评估框架,重现AI研究工作负载,从可用性和性能维度评估POSIX存储系统,通过案例比较Lustre和NFS存储系统,为不同环境选合适方案并辅助集群设计。
AI 中文摘要
人工智能研究的快速发展得益于对GPU集群的大量投资,但存储系统在实现高效研究工作流程中的关键作用常被忽视。与传统HPC工作负载不同,人工智能研究优先考虑研究人员的生产力和迭代的便利性。从业者在扩展到专用存储后端之前,依赖符合POSIX的文件系统进行无缝原型设计、调试和实验。因此,主要选择标准不仅是峰值吞吐量,还包括在POSIX兼容、对研究人员友好的接口内的性能。然而,现有基准仅根据峰值性能评估存储系统,无法捕捉现实世界人工智能研究中突发的、异构的IO模式。我们引入了PRISM评估框架,该框架重现了代表性的人工智能研究工作负载,包括数据摄取、检查点IO和开发人员工作流程,以在GPU集群上从可用性和性能维度评估和鉴定POSIX存储系统。使用PRISM,我们能够在多个研究工作负载维度上比较基于Lustre和NFS的POSIX存储系统,并为不同环境选择合适的存储解决方案。作为我们环境中的一个具体案例研究,我们观察到在分布式检查点加载用例中,基于闪存的NFS解决方案比基于闪存的Lustre解决方案性能高出3倍,这有助于我们做出明智的集群设计。
英文摘要
Large GPU clusters for AI research rely on POSIX-based storage spanning code authoring, data preparation, and model checkpointing. The metric that matters is researcher iteration speed, not just peak throughput - but current benchmarks target the latter: synthetic tools stress peak bandwidth and IOPS, emulators replay captured training I/O, and neither execute real framework operations or captures the bursty, heterogeneous, metadata-heavy patterns of research. We present PRISM, an extensible benchmark suite that executes real workloads across key stages of AI research including version control, environment setup, data preparation, data loading, model checkpointing, and synthetic data generation. PRISM exercises what researchers actually run and thereby serves not only as a benchmark but also as a qualification gate for new storage systems before they are consumed by researchers. We share our experience using PRISM for over 18 months in AI research clusters with thousands of Nvidia hopper class GPUs consuming petabyte scale Lustre and NFS based storage systems in characterizing research workloads, identifying the appropriate system for different cluster environments and detecting performance regressions. As an example, PRISM surfaced an openat lock contention pathology that caused an 8x slowdown under Fully Sharded Data Parallel (FSDP) checkpointing which went uncaught in emulated benchmarks. We plan to open source PRISM as a practical tool for selecting and validating storage for AI research infrastructure.