AI 中文总结
研究针对Linux大缓冲区读取问题,提出MARS多阶段加速读取堆栈,按数据结构和依赖性分阶段操作,提前处理页面错误和复制数据,在多种测试中提升带宽,加速查询和模型加载。
AI 中文摘要
大缓冲区读取越来越多地将数据密集型应用程序与高速存储连接起来。Linux 缓冲读取主要利用了其中一个好处,而在大读取中,其传统的交错路径会放大元数据和串行编排开销。我们提出了 MARS,一种用于同步大缓冲区缓冲读取的多阶段加速读取堆栈。MARS 将每个大范围读取视为一个工作单元,并按数据结构和依赖性分阶段进行页面缓存操作。在 I/O 等待期间,它处理用户缓冲区页面错误并提前执行可重新排序的数据复制。机会主义内核工作线程然后并行复制剩余数据,并在后端提供足够并行性时,可选地并行提交 I/O。我们在 Linux 6.6.58 中实现了 MARS。对于 MiB 规模的 fio 读取,MARS 比 Linux 的带宽提高了 6.56 倍。在 RAID0 中的五个 NVMe SSD 上,128 MiB 随机读取时达到 36.87 GiB/s,是 Linux 的 4.44 倍。MARS 还将 DuckDB/Parquet 查询加速了 1.80 - 2.15 倍,将 ExecuTorch 模型加载加速了 3.17 - 3.61 倍。
英文摘要
Large-buffer reads increasingly connect data-intensive applications to high-speed storage. They amortize system-call overhead and create a larger in-kernel window for organizing page-cache work and submitting I/O. However, Linux buffered read primarily exploits only the former benefit. Within a large read, its conventional interleaved path repeatedly switches among fine-grained page-cache operations, amplifying metadata and serial orchestration overheads and failing to consistently expose enough in-flight requests to modern parallel SSDs. We present MARS, a multi-stage accelerated read stack for synchronous large-buffer buffered reads. MARS treats each large-range read as one unit of work and stages page-cache operations by data structure and dependency. During I/O waits, it handles user-buffer page faults and performs reorderable data copies early. Opportunistic kernel workers then copy remaining data in parallel and, when the backend provides sufficient parallelism, optionally submit I/O in parallel. We implement MARS in Linux 6.6.58. For MiB-scale fio reads, MARS improves bandwidth by up to 6.56 times over Linux. On five NVMe SSDs in RAID0, it reaches 36.87 GiB/s for 128 MiB random reads, 4.44 times Linux. MARS also accelerates DuckDB/Parquet queries by 1.80--2.15 times and ExecuTorch model loading by 3.17--3.61 times.
CommentsWe found that the evaluation used Linux's default read-ahead setting. Enlarging the read-ahead window lets Linux issue substantially more I/O before waiting and removes most of MARS's reported advantage in the evaluated cold-read workloads. This materially changes the paper's central performance interpretation, so we withdraw it