arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

使用 NIO Bench 对机器学习工作负载的存储系统进行基准测试

Benchmarking Storage Systems for Machine Learning Workloads Using NIO Bench

Jonathan W. Morris, Ionut Mistreanu, Connor Louie

arXiv 2609.05418首次发表:更新:

发表机构

University of California, Santa Cruz(加州大学圣克鲁兹分校)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出NIO Bench框架,通过两层追踪方法刻画六种ML模型的存储访问模式,发现I/O集中于数据准备、模型加载和检查点,且存储访问呈极端幂律,建议优化预取、页缓存固定和突发写入处理。

AI 中文摘要

机器学习训练工作负载对存储系统提出了独特的要求,然而大多数现有基准测试侧重于计算吞吐量,而非文件系统 I/O 行为。我们提出了一个基准测试框架——神经 I/O 基准(NIO Bench),该框架刻画了六种不同 ML 模型架构的存储访问模式:语言 Transformer、视觉 Transformer、扩散模型、脉冲神经网络、人工神经网络和强化学习。我们的框架采用两层追踪方法,将用于语义阶段上下文的 Python 级 I/O 钩子与用于完整系统调用覆盖(包括 DataLoader 工作子进程)的 Linux strace 相结合。我们在一个带有 Ceph 分布式文件系统的 Nautilus Kubernetes 集群上评估了全部六种模型。我们的结果表明,I/O 高度集中于数据准备、模型加载和模型检查点保存。我们还发现,一旦数据就绪,训练是计算密集型而非数据密集型,并且存储访问遵循极端幂律,即不到 10% 的文件占据了超过 90% 的传输字节数,而分布式存储上缓存未命中的读取尾延迟是主要的存储瓶颈。这些发现表明,为 ML 优化的存储系统应优先考虑激进的数据预取、页缓存固定以及对突发检查点写入的高效处理。

英文摘要

Machine learning training workloads place unique demands on storage systems, yet most existing benchmarks focus on computational throughput rather than file system I/O behavior. We present a benchmarking framework, Neural I/O Benchmark (NIO Bench), that characterizes storage access patterns across six diverse ML model architectures: Language Transformers, Vision Transformers, Diffusion Models, Spiking Neural Networks, Artificial Neural Networks, and Reinforcement Learning. Our framework employs a two-layer tracing approach combining Python-level I/O hooks for semantic phase context with Linux strace for complete syscall coverage including DataLoader worker subprocesses. We evaluate all six models on a Nautilus Kubernetes cluster with Ceph distributed file system. Our results reveal that I/O is heavily concentrated in data preparation, model loading, and model checkpointing. We also found that training is compute-bound rather than data-bound once data is staged, and that storage access follows an extreme power law where fewer than 10% of files account for over 90% of bytes transferred, and that read tail latency from cache misses on distributed storage is the primary storage bottleneck. These findings suggest that storage systems optimized for ML should prioritize aggressive data prefetching, page cache pinning, and efficient handling of bursty checkpoint writes.

Comments7 pages, 6 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑